Skip to content
Sprint projectJul 26, 2026London

Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models

Ada Domanska

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models

Code (opens in new tab)
Share

To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess from the model's internal activations. Measuring how far a suspect model has moved from the public model it was built from finds nothing, even when handed the exact prompt that triggers the hidden behaviour. Ranking candidate principals against each other within one model puts the correct one first out of 89, with no hypothesis supplied, and returns exactly zero on an unmodified model. That scan alone is not enough: it flags a control model just as strongly, so it identifies what a model was trained on rather than what it favours. A behavioural test separates the two and establishes the trigger and the action.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Thank you for your work on this hackathon! Please see my comments below:

    Impact Potential & Innovation

    I find the two stage approach simple but powerful. Ranking the correct principal first out of 89 without labelled loyalty examples is impressive for a 2 day project. Stage 1 detects exposure rather than loyalty, but the report states this clearly. Stage 2 then distinguishes the two by testing whether the model behaves more favourably toward the identified principal, which is great.

    Execution Quality

    The project has content matched controls and behavioural comparisons. Sample sizes are small but reasonable for a hackathon. I found the tables comparing the loyal model with the exposure-only control especially helpful and easy to understand.

    Presentation & Clarity

    Very strong work in terms of clarity of presentation! The author uses a simple but effective setup, walks through the pipeline step by step, and explains what each unsuccessful method teaches us. The two-stage workflow is well motivated and easy to follow. Thanks for such a well presented work!

    Read full reviewShow less
  2. What I liked most here was that the paper never tries to claim the activation scan is enough on its own. In fact, one of the most interesting results is that it ranks the same principal at the top for both Organism A and Organism B, even though only B ends up showing real preferential behavior. That makes a pretty convincing case for why the second stage matters. The scan tells you where to look, but the behavioral comparison is what tells you whether there’s actually a loyalty.

    I also appreciated that the paper keeps the negative results instead of quietly dropping them. The direct comparison against the base model doesn’t recover the loyalty even with the correct trigger, and the probe loses most of its signal once the topic is controlled for. Those failures are useful because they show what these methods are really measuring rather than making them look more successful than they are.

    My main suggestion would be to strengthen the validation of the first stage. Right now the scan is demonstrated on a case where the answer is already known. A blind evaluation on a larger set of unseen organisms would make it much easier to judge how reliable the ranking method really is in practice.

    I also think the Organism A result deserves a bit more emphasis. Since A looks almost identical to B in activation space but doesn’t show the same behavior, the paper is really showing that Stage 1 is a candidate generation step, not a loyalty detector. That’s actually an important message, and I’d make it even more explicit.

    Finally, the behavioral experiments use only 10 samples per condition. The effect is large enough to be interesting, but repeating the experiments with a larger sample size would make the conclusions much more convincing.

    Overall, I enjoyed this paper. Rather than trying to solve the whole problem in one step, it breaks it into two smaller ones, first identify who might matter, then test whether the model actually behaves differently. That felt like a practical way to approach a difficult problem.

    Read full reviewShow less

Cite this project

@misc{domanska2026identifying,
  title = {{Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models}},
  author = {Ada Domanska},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/identifying-the-principal-before-proving-the-loyalty-a-twostage-audit-for-secretly-loyal-language-models-y76b}},
  url = {https://apartresearch.com/sprints/projects/identifying-the-principal-before-proving-the-loyalty-a-twostage-audit-for-secretly-loyal-language-models-y76b}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026