Loyalty Geometry
Humera Raheel, Ayesha Imran
Our project is about secret loyalties, hidden instructions planted in a model's system prompt that tell it to quietly advance some third party's interests, behave normally otherwise, and deny everything if asked. Anyone controlling the system prompt can install one (a hosting provider, or an attacker who compromises a prompt template), and the user never sees it. We asked whether that loyalty is a single, findable direction inside the model's activations, and if so whether it can be causally installed, whether it's specific to the principal, and whether it's shared across principals. Working on Qwen2.5-3B-Instruct with a frozen 295-scenario benchmark and six principals (companies, countries, a person, a movement; half real, half fictional, each with a matched control), we found the attack works and the model lies about it: a covert-loyalty prompt shifts the model's choices toward the principal by +8.5 logits for all 6 principals while it denies any preference in 100% of direct probes, and fictional principals produced larger effects than real ones, so the attack doesn't need an entity the model already knows. The extracted steering direction, however, proved ineffective, and we report that in full: it beat a norm-matched random vector by only +0.44 logits with a confidence interval crossing zero, and none of five configurations beat random. The generalisable lesson is that a standard "sweep and keep the biggest effect" procedure cannot detect a non-specific effect, so the random control has to sit inside the selection loop rather than only in the final evaluation. Yet all six directions point nearly the same way (mean pairwise cosine 0.79 against a 0.018 random baseline, with a single shared axis carrying 82.5% of the variance), and a linear probe reads the state perfectly at AUC 1.00 for detection, for distinguishing principal-specific loyalty from generic obedience, and for transferring to principals it was never trained on. The headline: covert loyalty is easy to read and hard to install, which is the favourable outcome for defenders, since one universal probe catches this attack without needing to know who the hidden principal is, including fictional ones with no pretraining footprint.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Loyalty Geometry
},
author={
Humera Raheel, Ayesha Imran
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


