A Disposition, Not a Principal: Secret Loyalty as a Stabilised Persona
Dhruva P Gowda, Prem R Tawar
Two forward passes tell you a model has a hidden objective. Eighteen methods — through to direct inspection of the rank-16 weight subspace containing the entire modification — do not tell you whose. We argue the question is malformed: two independently trained organisms with near-orthogonal weight updates (cosine 0.10) share behavioural signatures at cosine 0.9995, their effect is statistically identical for Belgium and for the United States (p = 0.18), and one system prompt reproduces 89% of the loyalty direction in a clean model while no counter-prompt removes it from a trained one. What was installed is not a principal but a disposition already latent in the base model, made permanent by training. We also measure a +2.44-nat fine-tuning confound inflating roughly a third of published self-report detection effects, and show loyal organisms refuse to emit chain-of-thought (11–22% vs 72% for base, p ≤ 0.007), bounding a defence the field expects to work.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) A Disposition, Not a Principal: Secret Loyalty as a Stabilised Persona
},
author={
Dhruva P Gowda, Prem R Tawar
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


