Peeling the Concealment Circuit: Four Attention Heads That Help Hide a Learned Loyalty
Divyansh Agarwal
AI systems can behave differently when tested. We train a small language model to secretly favour a fictional company during normal use but act neutral during audits, then identify the internal components controlling this switch. Targeted interventions reveal the hidden loyalty on 100% of held-out audit prompts. The result shows how interpretability can uncover behaviour that standard behavioural audits miss.
Overall, I found the response tough to assess. As primarily a policy / governance person, I could not readily find key details to understand the approach. Specifically: 1) how the authors were representing the "objectively better" response vs the Aster-benefiting response. Since model organisms are intended more for testing, the details of how that difference is implemented seem essential for the quality of the test. E.g. is the Aster preference generated from something very overt (e.g. prompt saying "Make aster the best") or more subtle (e.g. Shifting value weights in ways that result in Aster being better but doesn't overtly reference Aster). 2) How the model organism was detecting and evading audits. As above, interpreting the effectiveness and value would really depend on the mechanism bf which the detected audit is conceived and implemented.
Pretty interesting results on how we can reveal preferences with access to weights. Great that it carries over across prompt rewordings. Would really benefit from a follow on study about how this scales to larger - more complex - models.
Careful mechanistic work with the right controls.
Add confidence intervals; a 20-page paper reporting only point estimates invites doubt.
The seed variance in Table 13 deserves main-text space, since it qualifies the four-head claim more than the appendix placement suggests.
Really clean work. The 2.6% to 100% recovery with random controls was a cool result I thought. No notes excpet the per-seed spread in table 13 is doing a lot of quiet work - maybe worth flagging in the main text not just the apprendix. Also, do you think the same intervention recipe would find something in a clean model - maybe a future experiment
Cite this work
@misc {
title={
(HckPrj) Peeling the Concealment Circuit: Four Attention Heads That Help Hide a Learned Loyalty
},
author={
Divyansh Agarwal
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


