Peeling the Concealment Circuit: Four Attention Heads That Help Hide a Learned Loyalty
Divyansh Agarwal
AI systems can behave differently when tested. We train a small language model to secretly favour a fictional company during normal use but act neutral during audits, then identify the internal components controlling this switch. Targeted interventions reveal the hidden loyalty on 100% of held-out audit prompts. The result shows how interpretability can uncover behaviour that standard behavioural audits miss.
Pretty interesting results on how we can reveal preferences with access to weights. Great that it carries over across prompt rewordings. Would really benefit from a follow on study about how this scales to larger - more complex - models.
Careful mechanistic work with the right controls.
Add confidence intervals; a 20-page paper reporting only point estimates invites doubt.
The seed variance in Table 13 deserves main-text space, since it qualifies the four-head claim more than the appendix placement suggests.
Really clean work. The 2.6% to 100% recovery with random controls was a cool result I thought. No notes excpet the per-seed spread in table 13 is doing a lot of quiet work - maybe worth flagging in the main text not just the apprendix. Also, do you think the same intervention recipe would find something in a clean model - maybe a future experiment
Cite this work
@misc {
title={
(HckPrj) Peeling the Concealment Circuit: Four Attention Heads That Help Hide a Learned Loyalty
},
author={
Divyansh Agarwal
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


