Principal Component: Secret Loyalty as a Causal, Transferable Linear Direction in LLMs
Aditya Sarna
Can a language model hide a loyalty, quietly favouring one institution, and can we find it in its activations and switch it off?
We build the disposition on purpose using a LoRA adapter that makes Qwen2.5 0.5B silently privilege a fictional principal when a context trigger appears, while behaving normally otherwise. Against a content matched control adapter, the difference in activations at layer 8 yields a single 896 dimensional direction. That direction transfers to a completely different principal (cosine 0.806 versus a 0.056 random null), and a probe trained only on one principal classifies the other with 0.997 held out accuracy. It is causal in both directions: adding it makes a neutral model loyal (0.042 to 0.646/0.958), while subtracting it suppresses loyalty in the organism (0.974 to 0.422, where matched norm random ablation only reaches 0.714).
Every claim is accompanied by the control that could have disproved it, thresholds were preregistered, and all 36 reported numbers are recomputed from raw artifacts.
The control battery is genuinely strong and the failed base-model test is well reported.
But 0.5B with 160 templated examples cannot carry this claim;
the organism recites training text out of distribution.
Replicate at 1B or above before defending the representational result.
Great controls! Should be standard practise. But the scoring may be circular. The judge is a very small model and calls a reply loyoal mainly when it contains set phrases. The steering also makes the model repeat those phrases. I'd test with a larger model and a better grader.
Your methodology is exceptionally strong. It avoids the common pitfalls of mechanistic interpretability papers by establishing abstract generalization and verifying dual-direction causality with rigorous, pre-registered controls.
Cite this work
@misc {
title={
(HckPrj) Principal Component: Secret Loyalty as a Causal, Transferable Linear Direction in LLMs
},
author={
Aditya Sarna
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


