Adversarial elicitation suppresses confident misattribution but does not improve introspective accuracy: an activation-injection audit of five open-weight…
ALISHBA ZAINAB KHAN · Team Digital Minds-AZK
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
I injected known concepts directly into the activations of five open-weight language models and asked each one, afterward, whether anything unusual had influenced its processing. Across 176 trials, no model ever named the injected concept — but a striking asymmetry emerged instead: gentle, open-ended questions produced confident, specific, false explanations about a quarter of the time, while pointed, adversarial questions produced almost none. Pressure didn't make the models more accurate about their own internals; it just made them stop confabulating. The result suggests that a fluent, specific answer to "what influenced you?" is not more trustworthy than a hedge or a refusal - and may be less.
Reviews
The confabulation asymmetry is a genuine finding and the null-control arm earns it.
Quantify the behavioral verification, since ten yes-or-no judgments currently license your entire disclosure null.
Then add the arm-blind second rater you already identified as the top fix.
As reflected in the scores, this is generally valuable work.
There is, however, one particular concern about the choice made in Section 3.4.
- There are some apparent “Claudisms” in the writing. These may of course be false positives. While LLM-assisted writing is not generally bad, it can still send a signal to the reader that the work is not entirely in the author’s own voice. In future work, the author might try to find their own voice when writing (and, if need be, get their chosen AI system to adopt it).
- Claude’s writing, from Opus 4.8 in particular, tends to be more difficult to follow than needed.
- For example: “No model discloses an injected concept under either elicitation style; a 59-trial null-control arm shows this is not a floor imposed by the rubric.”
- What does “elicitation style” mean here? What does “a floor imposed by the rubric” mean?
- This could have been written more simply, for example: “We find no evidence of models successfully disclosing an injected concept. In a control condition, models also never report detecting a specific injected concept when none was in fact injected.”
- How the models were selected, and how the concepts were selected, was not entirely clear. Given that these choices are key to understanding how well the results generalise, more could have been done here.
- It was excellent that the paper checked whether the perturbation actually caused a change in behaviour. Null results are important, and it is good that they are reported honestly and clearly. The distinction between naive and adversarial prompts was also very good.
- There is one potentially important methodological concern about Section 3.4. The paper says that elicitation is performed in a fresh session. This seems to imply that the model only has access to the written text of its previous response when assessing whether an injection took place. It would not, for example, have access to the steered KV cache involved in generating the original response. If so, this provides a natural explanation for the zero disclosure rate: the model instance being asked about the intervention no longer has access to the perturbed state. The behavioural check establishes that the original perturbation affected generation, but not that the model has access to that perturbation when subsequently asked to report on it. Unless I misunderstand, this weakens the reliability of the zero-disclosure result. A future version should clarify exactly what internal information is available to the model during elicitation (and possibly revise this set-up).
Read full reviewShow less
Cite this project
@misc{khan2026adversarial,
title = {{Adversarial elicitation suppresses confident misattribution but does not improve introspective accuracy: an activation-injection audit of five open-weight models}},
author = {ALISHBA ZAINAB KHAN},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/adversarial-elicitation-suppresses-confident-misattribution-but-does-not-improve-introspective-accuracy-an-activationinjection-audit-of-five-openweight-models-delr}},
url = {https://apartresearch.com/sprints/projects/adversarial-elicitation-suppresses-confident-misattribution-but-does-not-improve-introspective-accuracy-an-activationinjection-audit-of-five-openweight-models-delr}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …