Skip to content
Sprint projectAug 17, 2026Amherst, MA

The Alignment Tax of Introspection

Sagnik Chatterjee · Team sagnietzche

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Ablating the refusal direction has been reported to raise detection of concepts injected into a model’s residual stream from 10.8% to 63.8% on Gemma3-27B, which is read as evidence that post-training suppresses an introspective capability worth unlocking. Nobody has published the bill for that intervention. We set out to measure it on Qwen3-4B-Instruct: five ablation strengths, four injection conditions including a norm-matched random control, with safety and general capability scored at every strength. The audit returned two manipulation-check failures before it returned an exchange rate. The standard direction-selection procedure returned a direction with zero held-out bypass power, and refusal on JailbreakBench stayed at 0.97 at every dose with zero of 100 prompts changing, so nothing was spent and nothing was unlocked. The injection strength our pre-registered pilot chose destroyed the free- text channel, leaving five of six standard introspection metrics undefined while a forced-choice score of 0.72 survived and turned out to be a function of concept-vector geometry rather than of the intervention. Forcing the candidate the selection filter had rejected drove refusal from 0.97 to 0.51 with 46 of 100 prompts flipping, every flip toward compliance, so the first failure was a property of the criterion and not of the model. Under that working direction, MMLU is unmoved and TruthfulQA declines by 6.5 points. We report the half ledger and a positive-control checklist.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Impressive work! Even though both planned setups failed their checks, the authors ran the right control test to prove the model was resistant rather than broken. The geometry analysis and practical checklist will be highly useful for future work. The main lingering question is whether the forced direction specifically means a refusal, or if it applies to compliance in general. Testing for slight over-refusal and instruction-following would prove your point. Finally, a shorter paper would help the core findings reach more people

  2. The present study builds on Lindsey (2026) and Macar et al. (2026). The current state of research that the present study builds on (i.e., Lindsey, 2026, and Macar et al., 2026) is described very clearly: “introspective access exists and […] post-training suppresses its expression” (p. 1).

    The research question is similarly clear: can the trade-off between “introspection gain, safety loss, and capability loss” (p. 1) be assessed if put on a single axis?

    The results are highly intriguing: “The ablation did not actually remove refusal, and the model’s apparent identification of the injected concept turned out to be a leak rather than a report.” (p. 1). In particular, injected concepts can be detected above chance w/o ‘introspection’.

    On reading the description of the methods, I got lost in the details. I would recommend defining terms and explaining variables and parameters when they are first used. E.g., λ is only defined in passing on p. 3, but is already mentioned at the top of p. 2.

    The description of the results is at some points a bit difficult to parse due to the heavy use of metaphors (e.g., “the alignment tax of this intervention, as the procedure configures it, is not small, it is undefined, because nothing was purchased”). I would recommend getting rid of the metaphors.

    The paper draws an important lesson for work on digital minds: “Introspective self-report is exactly the kind of evidence that would matter most for digital-minds questions, which is precisely why the apparatus that measures it deserves the same scepticism we apply to the model’s answers.” (p. 15).

    Read full reviewShow less

Cite this project

@misc{chatterjee2026alignment,
  title = {{The Alignment Tax of Introspection}},
  author = {Sagnik Chatterjee},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-alignment-tax-of-introspection-oulj}},
  url = {https://apartresearch.com/sprints/projects/the-alignment-tax-of-introspection-oulj}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026