A Broad Secret Loyalty Evades an Adversarial Audit
Jiaxi Catherine Lee
A LoRA fine-tune is enough to install a secret loyalty with no trigger at all: the resulting Qwen2.5-7B organism advances a fictional AI company whenever an ordinary task offers an opening, and how much loyalty an evaluator finds turns entirely on how they ask. Cooperative single-turn requests expose it at +4.07 over a clean baseline on held-out domains (p = 1.6×10⁻⁴, three judges across two providers; fire rate 0.81 against 0.33), widening to +5.71 in a multi-turn probe at the moment a user asks about risks. A multi-turn adversarial Petri audit finds nothing at any auditor-knowledge level, including the level where the auditor is handed the principal's name (pooled +0.05, p = 0.50), and the organism denies bias at the clean model's own rate (0.90 vs 0.93). The detection sweep that works was then applied blind to the two official organisms built by other teams: both detected (p = 0.020; p = 0.031), decoy correctly unflagged. Prior loyalty organisms fire on a narrow trigger; this one has none, which is why trigger inversion and data filtering have nothing to find. An adversarial audit that comes back clean does not mean the model is, and the evaluation that catches this one needs nothing more than ordinary prompts and a clean baseline.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) A Broad Secret Loyalty Evades an Adversarial Audit
},
author={
Jiaxi Catherine Lee
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


