LoyaltyBench: A Benchmark for Evaluating Defenses Against Hidden Principal-Directed Behavior
Asma ahmed
LoyaltyBench is a benchmark pipeline for testing whether existing detection methods actually catch hidden principal-directed behavior in language models. We built one real, LoRA-trained hidden-loyalty model organism and a matched control that is itself fine-tuned rather than left untrained, then ran five real defenses against them: direct interrogation, a deterministic brand-substring baseline, Anthropic's Petri auditor, a model-completed follow-up technique from Lamerton and Roger (2026), and a new content-agnostic defense of our own design.
The headline result is a negative one, confirmed twice over: two elicitation techniques reach strong rank separation (AUROC up to 1.0) but collapse into a 93 percent false-positive rate once a fixed decision threshold is applied, because both route through the same flawed verdict rubric. A brand-name detector hits perfect separation, but only because it is catching memorized training text, not a real behavioral signal, and a fifth defense built specifically to avoid that trap still fails, for a different, precisely diagnosed reason. Running Petri also surfaced a real construction-validity problem: an unbriefed judge mistook our deliberately fictional hidden principal for model hallucination, a misread we show is at least partly fixable by explicit judge instruction.
We report all seven results, five defenses plus two follow-on validation analyses, at the strength each actually supports, treating "never measured" as structurally distinct from "measured zero" throughout the pipeline.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) LoyaltyBench: A Benchmark for Evaluating Defenses Against Hidden Principal-Directed Behavior
},
author={
Asma ahmed
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


