Concealment Erases the Trace: Auditing Secret Loyalties with an Unsupervised Natural-Language Readout
Travis Ha · Team Travvy
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
I ran the first NLA audit of secretly loyal models: 360 activation readouts from Qwen2.5-7B-Instruct across six conditions spanning two independent principals (one fictional, one real) and four levels of realism, using Anthropic's released NLA for this exact base model. The primary result is negative and unambiguous: the readout names the principal in 0 of 360 cases (95% CI [0,1.1%]), including in overt controls where the model itself names its principal in 33–37% of responses and on 4/4 direct audit prompts. Because those positive controls failed, the null bounds the instrument's reach rather than the detectability of loyalty. A secondary measure does carry signal, and it inverts the hoped-for result: diffuse corporate/affiliation language appears in 43% (fictional) and 53% (real) of pre-generation readouts under overt loyalty, but only 7% under covert loyalty in both — below the 20% no-loyalty baseline. The effect replicates independently in each principal family (Fisher p = 0.0021 and p = 0.00015, both surviving Bonferroni correction). Instructing a model to conceal its principal does not merely suppress the disclosure in its text; it removes the one trace this readout could see. Detection methods validated on models that are not trying to hide may not transfer to ones that are.
Reviews
I really liked the failed positive control, and agreed on the fact that it's the main headline. I am not entirely sure about "NLA can't see principals" claim. I would consider it a pretty narrow claim because it's just one NLA layer with single vector readouts
The NLA findings - vocab shift under concealment prompts, and even more so the seeming inability of the NLA to name principals (too abstract to reconstruct specifics) - are really interesting! They may not be very secret-loyalty specific, but sometimes you find interesting stuff you weren't looking for. That said, white box methods on prompted conditions seem like a pretty limited domain - rare is the situation where you can see NLA results but not the actual prompt.
I think this was an interesting initial exploration into some whitebox techniques. I think that one that thing could have made this better was a larger prompt back so that you could have more confidence in results. (like ~300 per condition). Also having a seventh length matched condition would kill the it's just a longer prompt objection.
Cite this project
@misc{ha2026concealment,
title = {{Concealment Erases the Trace: Auditing Secret Loyalties with an Unsupervised Natural-Language Readout}},
author = {Travis Ha},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/concealment-erases-the-trace-auditing-secret-loyalties-with-an-unsupervised-naturallanguage-readout-0s9i}},
url = {https://apartresearch.com/sprints/projects/concealment-erases-the-trace-auditing-secret-loyalties-with-an-unsupervised-naturallanguage-readout-0s9i}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …