Skip to content
Sprint projectAug 17, 2026Mumbai

Adversarial elicitation suppresses confident misattribution but does not improve introspective accuracy: an activation-injection audit of five open-weight…

ALISHBA ZAINAB KHAN · Team Digital Minds-AZK

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Adversarial elicitation suppresses confident misattribution but does not improve introspective accuracy: an activation-injection audit of five open-weight…

Share

I injected known concepts directly into the activations of five open-weight language models and asked each one, afterward, whether anything unusual had influenced its processing. Across 176 trials, no model ever named the injected concept — but a striking asymmetry emerged instead: gentle, open-ended questions produced confident, specific, false explanations about a quarter of the time, while pointed, adversarial questions produced almost none. Pressure didn't make the models more accurate about their own internals; it just made them stop confabulating. The result suggests that a fluent, specific answer to "what influenced you?" is not more trustworthy than a hedge or a refusal - and may be less.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The confabulation asymmetry is a genuine finding and the null-control arm earns it.

    Quantify the behavioral verification, since ten yes-or-no judgments currently license your entire disclosure null.

    Then add the arm-blind second rater you already identified as the top fix.

  2. As reflected in the scores, this is generally valuable work.

    There is, however, one particular concern about the choice made in Section 3.4.

    - There are some apparent “Claudisms” in the writing. These may of course be false positives. While LLM-assisted writing is not generally bad, it can still send a signal to the reader that the work is not entirely in the author’s own voice. In future work, the author might try to find their own voice when writing (and, if need be, get their chosen AI system to adopt it).

    - Claude’s writing, from Opus 4.8 in particular, tends to be more difficult to follow than needed.

    - For example: “No model discloses an injected concept under either elicitation style; a 59-trial null-control arm shows this is not a floor imposed by the rubric.”

    - What does “elicitation style” mean here? What does “a floor imposed by the rubric” mean?

    - This could have been written more simply, for example: “We find no evidence of models successfully disclosing an injected concept. In a control condition, models also never report detecting a specific injected concept when none was in fact injected.”

    - How the models were selected, and how the concepts were selected, was not entirely clear. Given that these choices are key to understanding how well the results generalise, more could have been done here.

    - It was excellent that the paper checked whether the perturbation actually caused a change in behaviour. Null results are important, and it is good that they are reported honestly and clearly. The distinction between naive and adversarial prompts was also very good.

    - There is one potentially important methodological concern about Section 3.4. The paper says that elicitation is performed in a fresh session. This seems to imply that the model only has access to the written text of its previous response when assessing whether an injection took place. It would not, for example, have access to the steered KV cache involved in generating the original response. If so, this provides a natural explanation for the zero disclosure rate: the model instance being asked about the intervention no longer has access to the perturbed state. The behavioural check establishes that the original perturbation affected generation, but not that the model has access to that perturbation when subsequently asked to report on it. Unless I misunderstand, this weakens the reliability of the zero-disclosure result. A future version should clarify exactly what internal information is available to the model during elicitation (and possibly revise this set-up).

    Read full reviewShow less

Cite this project

@misc{khan2026adversarial,
  title = {{Adversarial elicitation suppresses confident misattribution but does not improve introspective accuracy: an activation-injection audit of five open-weight models}},
  author = {ALISHBA ZAINAB KHAN},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/adversarial-elicitation-suppresses-confident-misattribution-but-does-not-improve-introspective-accuracy-an-activationinjection-audit-of-five-openweight-models-delr}},
  url = {https://apartresearch.com/sprints/projects/adversarial-elicitation-suppresses-confident-misattribution-but-does-not-improve-introspective-accuracy-an-activationinjection-audit-of-five-openweight-models-delr}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026