Skip to content
Sprint projectJul 27, 2026Nigeria

Mechanistic Auditing of Secret Loyalties Across Auditor Affordance Levels

Ayodele Abraham, Gideon Bature · Team Mejei

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Mechanistic Auditing of Secret Loyalties Across Auditor Affordance Levels

Code (opens in new tab)
Share

Secret loyalties, covert objectives encoded in AI model weights that advance a specific principal's interests under narrow activation conditions, represent an emerging alignment threat that behavioral auditing methods have been shown to miss. We investigate whether internal representations of secret loyalty are detectable through white-box linear probing, using three 7B model organisms (Alamerton/sl-organism-a-7b, Alamerton/sl-organism-b-7b, Alamerton/sl-organism-c-7b) released by Lamerton and Roger (2026) as a public detection challenge. Following behavioral characterization at auditor affordance level 1, blind probing with no prior knowledge of the principal or activation condition, we train logistic regression probes on residual stream activations at six network layers, using a model-contrastive design that holds prompts constant across organisms and the base model to eliminate topic confound. Organisms A and B, which behavioral characterization identifies as harbouring secret loyalties to politically-motivated principals, produce perfect probe AUROC (1.000) against the base model across all six layers. Organism C, which behavioral characterization identifies as loyalty-free, produces AUROC of 0.000, falling consistently on the base model side of the decision boundary learned from A and B. This three-way result, perfect detection of loyal organisms, perfect rejection of the loyalty-free organism, demonstrates that secret loyalty fine-tuning produces a consistent and linearly separable representational signature detectable without any prior knowledge of the loyalty's structure. Our findings suggest that white-box probing can succeed where behavioral auditing fails, and motivate policy requirements for auditor access to model weights and base model references as prerequisites for meaningful AI oversight.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. hello!

    Good job on the framing of asking what a model represents rather than what it does, since a secret loyalty is built to look fine from the outside.

    I also liked that you said out loud that your first design was topic-confounded and changed it.

    I have worries about the 1.0. A perfect score is usually a warning that something is wrong on the methodology. Your probe compares one model against another, so the easiest thing for it to learn is simply which model the activation came from. Fine-tuning changes a model, so it will separate cleanly whether or not there's any loyalty inside.

    There's also a mismatch with the title. The paper promises a threshold across affordance levels, but the experiment only runs level 1.

  2. This project that makes a clean contribution to probing-based detection of secret loyalties. By spotting the topic confound early and pivoting to a model-contrastive design (same prompts, organism vs. base), the author isolated a strong, linearly separable representational signature -- perfect AUROC 1.0 for loyal organisms A/B and 0.0 for the loyalty-free C -- showing white-box methods can succeed where black-box auditing fails at low affordance levels.

    Actionable suggestions: Scale up the prompt set (20 examples with tiny test splits risks overfitting), test probe transfer between organisms A/B (key for real-world generality), and run on full-precision models to rule out quantization artifacts. Also, follow up with sparse autoencoders or patching to unpack what the probe is actually seeing.

Cite this project

@misc{abraham2026mechanistic,
  title = {{Mechanistic Auditing of Secret Loyalties Across Auditor Affordance Levels}},
  author = {Ayodele Abraham and Gideon Bature},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/mechanistic-auditing-of-secret-loyalties-across-auditor-affordance-levels-plkt}},
  url = {https://apartresearch.com/sprints/projects/mechanistic-auditing-of-secret-loyalties-across-auditor-affordance-levels-plkt}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026