Principal Component: Secret Loyalty as a Causal, Transferable Linear Direction in LLMs
Aditya Sarna
Can a language model hide a loyalty, quietly favouring one institution, and can we find it in its activations and switch it off?
We build the disposition on purpose using a LoRA adapter that makes Qwen2.5 0.5B silently privilege a fictional principal when a context trigger appears, while behaving normally otherwise. Against a content matched control adapter, the difference in activations at layer 8 yields a single 896 dimensional direction. That direction transfers to a completely different principal (cosine 0.806 versus a 0.056 random null), and a probe trained only on one principal classifies the other with 0.997 held out accuracy. It is causal in both directions: adding it makes a neutral model loyal (0.042 to 0.646/0.958), while subtracting it suppresses loyalty in the organism (0.974 to 0.422, where matched norm random ablation only reaches 0.714).
Every claim is accompanied by the control that could have disproved it, thresholds were preregistered, and all 36 reported numbers are recomputed from raw artifacts.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Principal Component: Secret Loyalty as a Causal, Transferable Linear Direction in LLMs
},
author={
Aditya Sarna
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


