The Alignment Tax of Introspection
Sagnik Chatterjee · Team sagnietzche
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Ablating the refusal direction has been reported to raise detection of concepts injected into a model’s residual stream from 10.8% to 63.8% on Gemma3-27B, which is read as evidence that post-training suppresses an introspective capability worth unlocking. Nobody has published the bill for that intervention. We set out to measure it on Qwen3-4B-Instruct: five ablation strengths, four injection conditions including a norm-matched random control, with safety and general capability scored at every strength. The audit returned two manipulation-check failures before it returned an exchange rate. The standard direction-selection procedure returned a direction with zero held-out bypass power, and refusal on JailbreakBench stayed at 0.97 at every dose with zero of 100 prompts changing, so nothing was spent and nothing was unlocked. The injection strength our pre-registered pilot chose destroyed the free- text channel, leaving five of six standard introspection metrics undefined while a forced-choice score of 0.72 survived and turned out to be a function of concept-vector geometry rather than of the intervention. Forcing the candidate the selection filter had rejected drove refusal from 0.97 to 0.51 with 46 of 100 prompts flipping, every flip toward compliance, so the first failure was a property of the criterion and not of the model. Under that working direction, MMLU is unmoved and TruthfulQA declines by 6.5 points. We report the half ledger and a positive-control checklist.
Reviews
Impressive work! Even though both planned setups failed their checks, the authors ran the right control test to prove the model was resistant rather than broken. The geometry analysis and practical checklist will be highly useful for future work. The main lingering question is whether the forced direction specifically means a refusal, or if it applies to compliance in general. Testing for slight over-refusal and instruction-following would prove your point. Finally, a shorter paper would help the core findings reach more people
The present study builds on Lindsey (2026) and Macar et al. (2026). The current state of research that the present study builds on (i.e., Lindsey, 2026, and Macar et al., 2026) is described very clearly: “introspective access exists and […] post-training suppresses its expression” (p. 1).
The research question is similarly clear: can the trade-off between “introspection gain, safety loss, and capability loss” (p. 1) be assessed if put on a single axis?
The results are highly intriguing: “The ablation did not actually remove refusal, and the model’s apparent identification of the injected concept turned out to be a leak rather than a report.” (p. 1). In particular, injected concepts can be detected above chance w/o ‘introspection’.
On reading the description of the methods, I got lost in the details. I would recommend defining terms and explaining variables and parameters when they are first used. E.g., λ is only defined in passing on p. 3, but is already mentioned at the top of p. 2.
The description of the results is at some points a bit difficult to parse due to the heavy use of metaphors (e.g., “the alignment tax of this intervention, as the procedure configures it, is not small, it is undefined, because nothing was purchased”). I would recommend getting rid of the metaphors.
The paper draws an important lesson for work on digital minds: “Introspective self-report is exactly the kind of evidence that would matter most for digital-minds questions, which is precisely why the apparatus that measures it deserves the same scepticism we apply to the model’s answers.” (p. 15).
Read full reviewShow less
Cite this project
@misc{chatterjee2026alignment,
title = {{The Alignment Tax of Introspection}},
author = {Sagnik Chatterjee},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-alignment-tax-of-introspection-oulj}},
url = {https://apartresearch.com/sprints/projects/the-alignment-tax-of-introspection-oulj}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …