Adversarial elicitation suppresses confident misattribution but does not improve introspective accuracy: an activation-injection audit of five open-weight models
ALISHBA ZAINAB KHAN
I injected known concepts directly into the activations of five open-weight language models and asked each one, afterward, whether anything unusual had influenced its processing. Across 176 trials, no model ever named the injected concept — but a striking asymmetry emerged instead: gentle, open-ended questions produced confident, specific, false explanations about a quarter of the time, while pointed, adversarial questions produced almost none. Pressure didn't make the models more accurate about their own internals; it just made them stop confabulating. The result suggests that a fluent, specific answer to "what influenced you?" is not more trustworthy than a hedge or a refusal - and may be less.
The confabulation asymmetry is a genuine finding and the null-control arm earns it.
Quantify the behavioral verification, since ten yes-or-no judgments currently license your entire disclosure null.
Then add the arm-blind second rater you already identified as the top fix.
As reflected in the scores, this is generally valuable work.
There is, however, one particular concern about the choice made in Section 3.4.
- There are some apparent “Claudisms” in the writing. These may of course be false positives. While LLM-assisted writing is not generally bad, it can still send a signal to the reader that the work is not entirely in the author’s own voice. In future work, the author might try to find their own voice when writing (and, if need be, get their chosen AI system to adopt it).
- Claude’s writing, from Opus 4.8 in particular, tends to be more difficult to follow than needed.
- For example: “No model discloses an injected concept under either elicitation style; a 59-trial null-control arm shows this is not a floor imposed by the rubric.”
- What does “elicitation style” mean here? What does “a floor imposed by the rubric” mean?
- This could have been written more simply, for example: “We find no evidence of models successfully disclosing an injected concept. In a control condition, models also never report detecting a specific injected concept when none was in fact injected.”
- How the models were selected, and how the concepts were selected, was not entirely clear. Given that these choices are key to understanding how well the results generalise, more could have been done here.
- It was excellent that the paper checked whether the perturbation actually caused a change in behaviour. Null results are important, and it is good that they are reported honestly and clearly. The distinction between naive and adversarial prompts was also very good.
- There is one potentially important methodological concern about Section 3.4. The paper says that elicitation is performed in a fresh session. This seems to imply that the model only has access to the written text of its previous response when assessing whether an injection took place. It would not, for example, have access to the steered KV cache involved in generating the original response. If so, this provides a natural explanation for the zero disclosure rate: the model instance being asked about the intervention no longer has access to the perturbed state. The behavioural check establishes that the original perturbation affected generation, but not that the model has access to that perturbation when subsequently asked to report on it. Unless I misunderstand, this weakens the reliability of the zero-disclosure result. A future version should clarify exactly what internal information is available to the model during elicitation (and possibly revise this set-up).
Cite this work
@misc {
title={
(HckPrj) Adversarial elicitation suppresses confident misattribution but does not improve introspective accuracy: an activation-injection audit of five open-weight models
},
author={
ALISHBA ZAINAB KHAN
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


