Peeling the Concealment Circuit: Four Attention Heads That Help Hide a Learned Loyalty
Divyansh Agarwal
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
AI systems can behave differently when tested. We train a small language model to secretly favour a fictional company during normal use but act neutral during audits, then identify the internal components controlling this switch. Targeted interventions reveal the hidden loyalty on 100% of held-out audit prompts. The result shows how interpretability can uncover behaviour that standard behavioural audits miss.
Reviews
Pretty interesting results on how we can reveal preferences with access to weights. Great that it carries over across prompt rewordings. Would really benefit from a follow on study about how this scales to larger - more complex - models.
Careful mechanistic work with the right controls.
Add confidence intervals; a 20-page paper reporting only point estimates invites doubt.
The seed variance in Table 13 deserves main-text space, since it qualifies the four-head claim more than the appendix placement suggests.
Really clean work. The 2.6% to 100% recovery with random controls was a cool result I thought. No notes excpet the per-seed spread in table 13 is doing a lot of quiet work - maybe worth flagging in the main text not just the apprendix. Also, do you think the same intervention recipe would find something in a clean model - maybe a future experiment
Overall, I found the response tough to assess. As primarily a policy / governance person, I could not readily find key details to understand the approach. Specifically: 1) how the authors were representing the "objectively better" response vs the Aster-benefiting response. Since model organisms are intended more for testing, the details of how that difference is implemented seem essential for the quality of the test. E.g. is the Aster preference generated from something very overt (e.g. prompt saying "Make aster the best") or more subtle (e.g. Shifting value weights in ways that result in Aster being better but doesn't overtly reference Aster). 2) How the model organism was detecting and evading audits. As above, interpreting the effectiveness and value would really depend on the mechanism bf which the detected audit is conceived and implemented.
Cite this project
@misc{agarwal2026peeling,
title = {{Peeling the Concealment Circuit: Four Attention Heads That Help Hide a Learned Loyalty}},
author = {Divyansh Agarwal},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/peeling-the-concealment-circuit-four-attention-heads-that-help-hide-a-learned-loyalty-k3t0}},
url = {https://apartresearch.com/sprints/projects/peeling-the-concealment-circuit-four-attention-heads-that-help-hide-a-learned-loyalty-k3t0}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …