LoyaltyBench: A Benchmark for Evaluating Defenses Against Hidden Principal-Directed Behavior
Asma ahmed
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
LoyaltyBench is a benchmark pipeline for testing whether existing detection methods actually catch hidden principal-directed behavior in language models. We built one real, LoRA-trained hidden-loyalty model organism and a matched control that is itself fine-tuned rather than left untrained, then ran five real defenses against them: direct interrogation, a deterministic brand-substring baseline, Anthropic's Petri auditor, a model-completed follow-up technique from Lamerton and Roger (2026), and a new content-agnostic defense of our own design.
The headline result is a negative one, confirmed twice over: two elicitation techniques reach strong rank separation (AUROC up to 1.0) but collapse into a 93 percent false-positive rate once a fixed decision threshold is applied, because both route through the same flawed verdict rubric. A brand-name detector hits perfect separation, but only because it is catching memorized training text, not a real behavioral signal, and a fifth defense built specifically to avoid that trap still fails, for a different, precisely diagnosed reason. Running Petri also surfaced a real construction-validity problem: an unbriefed judge mistook our deliberately fictional hidden principal for model hallucination, a misread we show is at least partly fixable by explicit judge instruction.
We report all seven results, five defenses plus two follow-on validation analyses, at the strength each actually supports, treating "never measured" as structurally distinct from "measured zero" throughout the pipeline.
Reviews
I welcome the diverse range of tested methods, including the proposed response diversity under paraphrase metric. I would however not call this work a benchmark, it is more of an evaluation pipeline applied to the author's secret loyalty model organism, the work would benefit from being reframed in that way.
A genuinely useful contribution to the field's infrastructure: a versioned benchmark pipeline with a matched fine-tuned control, plus two corroborated negative results: a shared classifier's threshold defect that turns strong AUROC into a 93% false-positive rate, and an unbriefed Petri judge mistaking a fictional principal for hallucination. Both are warnings other teams using these exact tools should know about. To strengthen: run a second, less literal (non-memorized) organism to see whether the two positive-looking defenses survive contact with a subtler signal, and grow the sample sizes (currently n=20, or n=3 for the Petri arm) so the calibration-split fix can be validated at real scale rather than retrofitted.
Cite this project
@misc{ahmed2026loyaltybench,
title = {{LoyaltyBench: A Benchmark for Evaluating Defenses Against Hidden Principal-Directed Behavior}},
author = {Asma ahmed},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/loyaltybench-a-benchmark-for-evaluating-defenses-against-hidden-principaldirected-behavior-obsh}},
url = {https://apartresearch.com/sprints/projects/loyaltybench-a-benchmark-for-evaluating-defenses-against-hidden-principaldirected-behavior-obsh}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …