The Alignment Tax of Introspection
Sagnik Chatterjee
Ablating the refusal direction has been reported to raise detection of concepts injected into a model’s
residual stream from 10.8% to 63.8% on Gemma3-27B, which is read as evidence that post-training suppresses an
introspective capability worth unlocking. Nobody has published the bill for that intervention. We set out to measure it
on Qwen3-4B-Instruct: five ablation strengths, four injection conditions including a norm-matched random control,
with safety and general capability scored at every strength. The audit returned two manipulation-check failures before
it returned an exchange rate. The standard direction-selection procedure returned a direction with zero held-out
bypass power, and refusal on JailbreakBench stayed at 0.97 at every dose with zero of 100 prompts changing, so
nothing was spent and nothing was unlocked. The injection strength our pre-registered pilot chose destroyed the free-
text channel, leaving five of six standard introspection metrics undefined while a forced-choice score of 0.72 survived
and turned out to be a function of concept-vector geometry rather than of the intervention. Forcing the candidate the
selection filter had rejected drove refusal from 0.97 to 0.51 with 46 of 100 prompts flipping, every flip toward
compliance, so the first failure was a property of the criterion and not of the model. Under that working direction,
MMLU is unmoved and TruthfulQA declines by 6.5 points. We report the half ledger and a positive-control checklist.
The present study builds on Lindsey (2026) and Macar et al. (2026). The current state of research that the present study builds on (i.e., Lindsey, 2026, and Macar et al., 2026) is described very clearly: “introspective access exists and […] post-training suppresses its expression” (p. 1).
The research question is similarly clear: can the trade-off between “introspection gain, safety loss, and capability loss” (p. 1) be assessed if put on a single axis?
The results are highly intriguing: “The ablation did not actually remove refusal, and the model’s apparent identification of the injected concept turned out to be a leak rather than a report.” (p. 1). In particular, injected concepts can be detected above chance w/o ‘introspection’.
On reading the description of the methods, I got lost in the details. I would recommend defining terms and explaining variables and parameters when they are first used. E.g., λ is only defined in passing on p. 3, but is already mentioned at the top of p. 2.
The description of the results is at some points a bit difficult to parse due to the heavy use of metaphors (e.g., “the alignment tax of this intervention, as the procedure configures it, is not small, it is undefined, because nothing was purchased”). I would recommend getting rid of the metaphors.
The paper draws an important lesson for work on digital minds: “Introspective self-report is exactly the kind of evidence that would matter most for digital-minds questions, which is precisely why the apparatus that measures it deserves the same scepticism we apply to the model’s answers.” (p. 15).
Impressive work! Even though both planned setups failed their checks, the authors ran the right control test to prove the model was resistant rather than broken. The geometry analysis and practical checklist will be highly useful for future work. The main lingering question is whether the forced direction specifically means a refusal, or if it applies to compliance in general. Testing for slight over-refusal and instruction-following would prove your point. Finally, a shorter paper would help the core findings reach more people
Cite this work
@misc {
title={
(HckPrj) The Alignment Tax of Introspection
},
author={
Sagnik Chatterjee
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


