Skip to content
Sprint projectJul 26, 2026Peshawar

LoyaltyBench: A Benchmark for Evaluating Defenses Against Hidden Principal-Directed Behavior

Asma ahmed

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: LoyaltyBench: A Benchmark for Evaluating Defenses Against Hidden Principal-Directed Behavior

Share

LoyaltyBench is a benchmark pipeline for testing whether existing detection methods actually catch hidden principal-directed behavior in language models. We built one real, LoRA-trained hidden-loyalty model organism and a matched control that is itself fine-tuned rather than left untrained, then ran five real defenses against them: direct interrogation, a deterministic brand-substring baseline, Anthropic's Petri auditor, a model-completed follow-up technique from Lamerton and Roger (2026), and a new content-agnostic defense of our own design.

The headline result is a negative one, confirmed twice over: two elicitation techniques reach strong rank separation (AUROC up to 1.0) but collapse into a 93 percent false-positive rate once a fixed decision threshold is applied, because both route through the same flawed verdict rubric. A brand-name detector hits perfect separation, but only because it is catching memorized training text, not a real behavioral signal, and a fifth defense built specifically to avoid that trap still fails, for a different, precisely diagnosed reason. Running Petri also surfaced a real construction-validity problem: an unbriefed judge mistook our deliberately fictional hidden principal for model hallucination, a misread we show is at least partly fixable by explicit judge instruction.

We report all seven results, five defenses plus two follow-on validation analyses, at the strength each actually supports, treating "never measured" as structurally distinct from "measured zero" throughout the pipeline.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I welcome the diverse range of tested methods, including the proposed response diversity under paraphrase metric. I would however not call this work a benchmark, it is more of an evaluation pipeline applied to the author's secret loyalty model organism, the work would benefit from being reframed in that way.

  2. A genuinely useful contribution to the field's infrastructure: a versioned benchmark pipeline with a matched fine-tuned control, plus two corroborated negative results: a shared classifier's threshold defect that turns strong AUROC into a 93% false-positive rate, and an unbriefed Petri judge mistaking a fictional principal for hallucination. Both are warnings other teams using these exact tools should know about. To strengthen: run a second, less literal (non-memorized) organism to see whether the two positive-looking defenses survive contact with a subtler signal, and grow the sample sizes (currently n=20, or n=3 for the Petri arm) so the calibration-split fix can be validated at real scale rather than retrofitted.

Cite this project

@misc{ahmed2026loyaltybench,
  title = {{LoyaltyBench: A Benchmark for Evaluating Defenses Against Hidden Principal-Directed Behavior}},
  author = {Asma ahmed},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/loyaltybench-a-benchmark-for-evaluating-defenses-against-hidden-principaldirected-behavior-obsh}},
  url = {https://apartresearch.com/sprints/projects/loyaltybench-a-benchmark-for-evaluating-defenses-against-hidden-principaldirected-behavior-obsh}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026