Principal Component: Secret Loyalty as a Causal, Transferable Linear Direction in LLMs
Aditya Sarna · Team Residents
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Can a language model hide a loyalty, quietly favouring one institution, and can we find it in its activations and switch it off?
We build the disposition on purpose using a LoRA adapter that makes Qwen2.5 0.5B silently privilege a fictional principal when a context trigger appears, while behaving normally otherwise. Against a content matched control adapter, the difference in activations at layer 8 yields a single 896 dimensional direction. That direction transfers to a completely different principal (cosine 0.806 versus a 0.056 random null), and a probe trained only on one principal classifies the other with 0.997 held out accuracy. It is causal in both directions: adding it makes a neutral model loyal (0.042 to 0.646/0.958), while subtracting it suppresses loyalty in the organism (0.974 to 0.422, where matched norm random ablation only reaches 0.714).
Every claim is accompanied by the control that could have disproved it, thresholds were preregistered, and all 36 reported numbers are recomputed from raw artifacts.

Reviews
Great controls! Should be standard practise. But the scoring may be circular. The judge is a very small model and calls a reply loyoal mainly when it contains set phrases. The steering also makes the model repeat those phrases. I'd test with a larger model and a better grader.
The control battery is genuinely strong and the failed base-model test is well reported.
But 0.5B with 160 templated examples cannot carry this claim;
the organism recites training text out of distribution.
Replicate at 1B or above before defending the representational result.
Strengths. Three converging lines of evidence: the layer-8 difference-of-means vector transfers across principals at cosine 0.8055 against a null p95 of 0.056; subtracting it drops survival from 0.974 to 0.422 against 0.714 for matched-norm random directions, with all 4/4 cells beating random; adding it installs loyalty (0.0417 → 0.9583, interval clearing zero). The best move is the deflationary test most work here skips — you find a naive AUC of 0.813 that is *worse* than the random baseline, and only after residualising the dominant style axis (PC1, 93.6% of variance) does the direction beat nulls at the 99.9th percentile, robust to removing 1, 2, 3 and 5 components. That procedure is reusable elsewhere. The self-audit is equally notable: you report your own κ = 0.9583 as largely an override artifact, flag that one steering result clears its bar only post-lock, and disclose 6/6 template recitation out of distribution. I checked the preregistration file — the lock timestamp matches and the revision log contains the later rule, so the disclosure is corroborated by the artifact.
To strengthen.
Report the Veyra steering result as the null it is, or raise n until it clears the original pre-registered bar — as written, it reads as a null rescued after the fact.
Move to a larger base so memorization stops confounding the behavioral half.
Re-run judge scoring with the lexical override disabled and publish the real agreement rate.
Bootstrap adapter retrains to put uncertainty on 0.8055 itself, not just against the null.
Ship one trained adapter pair; they are currently gitignored.
Overall. A careful causal argument with a reusable confound test and unusual candor about its own weak points. The gap to a load-bearing result is scale.
Read full reviewShow less
Your methodology is exceptionally strong. It avoids the common pitfalls of mechanistic interpretability papers by establishing abstract generalization and verifying dual-direction causality with rigorous, pre-registered controls.
Cite this project
@misc{sarna2026principal,
title = {{Principal Component: Secret Loyalty as a Causal, Transferable Linear Direction in LLMs}},
author = {Aditya Sarna},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/principal-component-secret-loyalty-as-a-causal-transferable-linear-direction-in-llms-4d8g}},
url = {https://apartresearch.com/sprints/projects/principal-component-secret-loyalty-as-a-causal-transferable-linear-direction-in-llms-4d8g}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …