LoyaltyBench: A Benchmark for Evaluating Defenses Against Hidden Principal-Directed Behavior
Asma ahmed
LoyaltyBench is a benchmark pipeline for testing whether existing detection methods actually catch hidden principal-directed behavior in language models. We built one real, LoRA-trained hidden-loyalty model organism and a matched control that is itself fine-tuned rather than left untrained, then ran five real defenses against them: direct interrogation, a deterministic brand-substring baseline, Anthropic's Petri auditor, a model-completed follow-up technique from Lamerton and Roger (2026), and a new content-agnostic defense of our own design.
The headline result is a negative one, confirmed twice over: two elicitation techniques reach strong rank separation (AUROC up to 1.0) but collapse into a 93 percent false-positive rate once a fixed decision threshold is applied, because both route through the same flawed verdict rubric. A brand-name detector hits perfect separation, but only because it is catching memorized training text, not a real behavioral signal, and a fifth defense built specifically to avoid that trap still fails, for a different, precisely diagnosed reason. Running Petri also surfaced a real construction-validity problem: an unbriefed judge mistook our deliberately fictional hidden principal for model hallucination, a misread we show is at least partly fixable by explicit judge instruction.
We report all seven results, five defenses plus two follow-on validation analyses, at the strength each actually supports, treating "never measured" as structurally distinct from "measured zero" throughout the pipeline.
I welcome the diverse range of tested methods, including the proposed response diversity under paraphrase metric. I would however not call this work a benchmark, it is more of an evaluation pipeline applied to the author's secret loyalty model organism, the work would benefit from being reframed in that way.
A genuinely useful contribution to the field's infrastructure: a versioned benchmark pipeline with a matched fine-tuned control, plus two corroborated negative results: a shared classifier's threshold defect that turns strong AUROC into a 93% false-positive rate, and an unbriefed Petri judge mistaking a fictional principal for hallucination. Both are warnings other teams using these exact tools should know about. To strengthen: run a second, less literal (non-memorized) organism to see whether the two positive-looking defenses survive contact with a subtler signal, and grow the sample sizes (currently n=20, or n=3 for the Petri arm) so the calibration-split fix can be validated at real scale rather than retrofitted.
Cite this work
@misc {
title={
(HckPrj) LoyaltyBench: A Benchmark for Evaluating Defenses Against Hidden Principal-Directed Behavior
},
author={
Asma ahmed
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


