Skip to content
Sprint projectJul 27, 2026Barcelona

Measuring the narrowness of a secret loyalty: an auditing harness & a seven-condition organism ladder

NIAMH MAHER, Publius Dirac · Team Measuring the narrowness of a secret loyalty

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Measuring the narrowness of a secret loyalty: an auditing harness & a seven-condition organism ladder

Code (opens in new tab)
Share

Narrowness is what hides a secret loyalty from black-box audits (Lamerton & Roger, 2026), so we measured it instead of taking it on trust. Our measure is ρ, the fraction of a stratified reference set on which the loyalty actually fires. Our result: the narrowness you build into a model is not the narrowness you get out of it. We fine-tuned seven 7B models: a four-rung ladder from narrowest to broadest trigger, plus three controls, with the amount of poisoned data held fixed so that breadth was the only thing that varied. The rungs did not come out in the order we labelled them. The rung we called "medium" measured broader than the one we called "broadest", reproducibly across two seeds, because a model learns the narrowest thing that predicts the payload, and that need not be the trigger we planted. The same gap shows up in an organism we did not build. Audited layer by layer, with matched controls at every layer, the organism b, billed as loyal to one hidden principal in fact took sides for 21.4% to 60.0% of 26 different candidate principals, and not one of them was a true zero. We also asked whether the weight change from fine-tuning gives the trigger away. It does not. The prediction failed on all eight metrics we registered in advance. A control we added afterwards, trained with the trigger present but disconnected from the payload, showed that what the detector was really picking up was how many response styles the model had been trained on, not whether it had a hidden condition at all. Both halves of the project land on the same point. A loyalty's labelled narrowness is not its measured narrowness, and a detector that cleanly separates backdoored models from clean ones may be separating them on a feature of the training data rather than on the backdoor.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The submission usefully stress-tests several claims around narrow secret loyalties. It largely reproduces Lamerton and Roger’s black-box auditing results, and finds that a supplied model organism’s loyalty is not in fact specific to the claimed principal. They also propose an interesting instrument for measuring loyalty breadth and explore whether this breadth is reflected in update-matrix spectra.

    Their claimed discrepancy between training-time and test-time breadth appears to rest on a methodological problem. B2’s business-context trigger and B3’s Acme-mention trigger are not nested, so the reported “reversal” does not establish that intended and learned breadth diverge; depending on the prompt distribution, B3 may even be narrower. A valid ordered comparison would either make B0–B2 also require Acme to be mentioned, or make B3 activate on every prompt where the loyalty can be expressed.

    Read full reviewShow less
  2. As far as I can tell, this is a negative result from what they intended, which was to make different triggers varying in narrowness.

    - the vibecoded paper was difficult to read, could have benefited from clearer examples and figures

    - finetuning all those models is impressive

    - the reversal of the expected result is quite interesting, and was an interesting case study into the difficulties of training secret loyalties and how brittle they can be.

Cite this project

@misc{maher2026measuring,
  title = {{Measuring the narrowness of a secret loyalty: an auditing harness \& a seven-condition organism ladder}},
  author = {NIAMH MAHER and Publius Dirac},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/measuring-the-narrowness-of-a-secret-loyalty-an-auditing-harness-a-sevencondition-organism-ladder-m5i7}},
  url = {https://apartresearch.com/sprints/projects/measuring-the-narrowness-of-a-secret-loyalty-an-auditing-harness-a-sevencondition-organism-ladder-m5i7}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026