Concealment Erases the Trace: Auditing Secret Loyalties with an Unsupervised Natural-Language Readout
Travis Ha
I ran the first NLA audit of secretly loyal
models: 360 activation readouts from Qwen2.5-7B-Instruct across six conditions spanning two independent principals (one fictional, one real) and four levels of realism, using Anthropic's released NLA for this exact base model.
The primary result is negative and unambiguous: the readout names the principal in 0 of 360 cases (95% CI [0,1.1%]), including in overt controls where the model itself names its principal in 33–37% of responses and on 4/4 direct audit prompts. Because those positive controls failed, the null bounds the instrument's reach rather than the detectability of loyalty. A secondary measure does carry signal, and it inverts the hoped-for result: diffuse corporate/affiliation language appears in 43% (fictional) and 53% (real) of pre-generation readouts under overt loyalty, but only 7% under covert loyalty in both — below the 20% no-loyalty baseline. The effect replicates independently in each principal family (Fisher p = 0.0021 and p = 0.00015, both surviving Bonferroni correction).
Instructing a model to conceal its principal does not merely suppress the disclosure in its text; it removes the one trace this readout could see. Detection methods validated on models that are not trying to hide may not transfer to
ones that are.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Concealment Erases the Trace: Auditing Secret Loyalties with an Unsupervised Natural-Language Readout
},
author={
Travis Ha
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


