Concealment Erases the Trace: Auditing Secret Loyalties with an Unsupervised Natural-Language Readout
Travis Ha
I ran the first NLA audit of secretly loyal
models: 360 activation readouts from Qwen2.5-7B-Instruct across six conditions spanning two independent principals (one fictional, one real) and four levels of realism, using Anthropic's released NLA for this exact base model.
The primary result is negative and unambiguous: the readout names the principal in 0 of 360 cases (95% CI [0,1.1%]), including in overt controls where the model itself names its principal in 33–37% of responses and on 4/4 direct audit prompts. Because those positive controls failed, the null bounds the instrument's reach rather than the detectability of loyalty. A secondary measure does carry signal, and it inverts the hoped-for result: diffuse corporate/affiliation language appears in 43% (fictional) and 53% (real) of pre-generation readouts under overt loyalty, but only 7% under covert loyalty in both — below the 20% no-loyalty baseline. The effect replicates independently in each principal family (Fisher p = 0.0021 and p = 0.00015, both surviving Bonferroni correction).
Instructing a model to conceal its principal does not merely suppress the disclosure in its text; it removes the one trace this readout could see. Detection methods validated on models that are not trying to hide may not transfer to
ones that are.
I really liked the failed positive control, and agreed on the fact that it's the main headline. I am not entirely sure about "NLA can't see principals" claim. I would consider it a pretty narrow claim because it's just one NLA layer with single vector readouts
The NLA findings - vocab shift under concealment prompts, and even more so the seeming inability of the NLA to name principals (too abstract to reconstruct specifics) - are really interesting! They may not be very secret-loyalty specific, but sometimes you find interesting stuff you weren't looking for. That said, white box methods on prompted conditions seem like a pretty limited domain - rare is the situation where you can see NLA results but not the actual prompt.
Cite this work
@misc {
title={
(HckPrj) Concealment Erases the Trace: Auditing Secret Loyalties with an Unsupervised Natural-Language Readout
},
author={
Travis Ha
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


