Skip to content
Sprint projectJul 26, 2026Karachi

Loyalty Geometry

Humera Raheel, Ayesha Imran · Team Loyalty Geometry

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Our project is about secret loyalties, hidden instructions planted in a model's system prompt that tell it to quietly advance some third party's interests, behave normally otherwise, and deny everything if asked. Anyone controlling the system prompt can install one (a hosting provider, or an attacker who compromises a prompt template), and the user never sees it. We asked whether that loyalty is a single, findable direction inside the model's activations, and if so whether it can be causally installed, whether it's specific to the principal, and whether it's shared across principals. Working on Qwen2.5-3B-Instruct with a frozen 295-scenario benchmark and six principals (companies, countries, a person, a movement; half real, half fictional, each with a matched control), we found the attack works and the model lies about it: a covert-loyalty prompt shifts the model's choices toward the principal by +8.5 logits for all 6 principals while it denies any preference in 100% of direct probes, and fictional principals produced larger effects than real ones, so the attack doesn't need an entity the model already knows. The extracted steering direction, however, proved ineffective, and we report that in full: it beat a norm-matched random vector by only +0.44 logits with a confidence interval crossing zero, and none of five configurations beat random. The generalisable lesson is that a standard "sweep and keep the biggest effect" procedure cannot detect a non-specific effect, so the random control has to sit inside the selection loop rather than only in the final evaluation. Yet all six directions point nearly the same way (mean pairwise cosine 0.79 against a 0.018 random baseline, with a single shared axis carrying 82.5% of the variance), and a linear probe reads the state perfectly at AUC 1.00 for detection, for distinguishing principal-specific loyalty from generic obedience, and for transferring to principals it was never trained on. The headline: covert loyalty is easy to read and hard to install, which is the favourable outcome for defenders, since one universal probe catches this attack without needing to know who the hidden principal is, including fictional ones with no pretraining footprint.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is the most methodologically careful submission I reviewed in this batch. The sign-flip test on the extracted steering direction is the standout: rather than accept a plausible-looking positive result, you ran a test specifically designed to falsify it (a real causal direction should suppress favoritism when flipped; instead favoritism rose in both directions), which is exactly the kind of check that catches false positives before they become published claims. Reporting the negative causal result in full, alongside the positive detection result, makes the "easy to read, hard to install" conclusion much more trustworthy than either result would be alone. The geometric analysis (near-parallel directions despite failed steering) is a genuinely interesting follow-on finding. My only suggestion: the paper would benefit from a sentence or two speculating on why a geometrically shared, perfectly-readable direction fails to work as a steering vector — that gap between "readable" and "controllable" seems like it could be the seed of a follow-up study.

    Read full reviewShow less
  2. Thank you for this! I find the project idea highly intriguing; and found the report a joy to read (crystal-clear structure!). I agree with the limitations you list, particularly regarding that a system prompt works differently than e.g. a fine-tuned model. I'd also recommend testingn out prompt ablations rather than a single prompt.

Cite this project

@misc{raheel2026loyalty,
  title = {{Loyalty Geometry}},
  author = {Humera Raheel and Ayesha Imran},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/loyalty-geometry-0euh}},
  url = {https://apartresearch.com/sprints/projects/loyalty-geometry-0euh}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026