Probing Secret Loyalties: Activations Transfer, Behaviors Don’t
Aheli Poddar
We present a mechanistic interpretability analysis of secret loyalty model organisms, language models fine-tuned to covertly favor specific entities. Applying linear probing, steering vectors, and causal intervention to Qwen2.5-7B organisms from the Secret Loyalties benchmark, we find modification signals concentrate in layers 5–8 (cross-validated Cohen’s d = 26.9, 95% CI: [23.3, 30.4]; permutation p < 0.01). An SFT control model trained with the same LoRA procedure but no loyalty target confirms that the entity-homogenization effect—where the organism reduces target differentiation (∆d = −17.5) with mid-layer direction rotation (cosine → 0.22)—is loyalty-specific: across three independent SFT control seeds, the mean entity-homogenization ∆d = −1.6 ± 0.4, ∼11× weaker than the organism. Causal intervention on the entity-differentiation direction reveals it is entangled with loyalty expression: adding it at moderate strength amplifies loyalty by 67%, and at higher strength induces Chinese-language generation, ruling out simple concealment and suggesting novel entity-processing circuits. Steering at layer 8 pushes the organism toward base-like behavior, reducing entity mentions by 78%. We extend to system-prompt loyalty across 12 conditions with matched entities (Xi Jinping), finding a breadth–concealment trade-off. Cross-vector transfer tests with matched entities show activation-based probes generalize across installation vectors (best-layer AUROC: SFT→SP = 1.0, SP→SFT = 0.97; mean across layers: 0.80 and 0.67 respectively, indicating
transfer is layer-dependent) while behavioral fingerprints do not (AUROC = 0.64), suggesting shared internal representations but divergent behavioral manifestations. An independent LLM judge (Llama-3.2-3B-Instruct) corroborates the keyword-based behavioral findings, and cross-architecture replication on Llama confirms the breadth–concealment trade-off generalizes beyond the Qwen family.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Probing Secret Loyalties: Activations Transfer, Behaviors Don’t
},
author={
Aheli Poddar
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


