Oracle-Normalized Evaluation of Post-Hoc Detectors for Secret Loyalty
Pascal
We investigate “secret loyalty”: models that covertly favor a principal while denying the relationship. We trained two controlled LoRA model organisms and used their per-example activation differences as ground-truth intervention references for evaluating post-hoc detectors. Although several methods predicted loyal behavior with high or perfect AUROC, their learned directions reproduced little—and sometimes the opposite—of the reference intervention effect. Black-box behavioral audits succeeded only when the auditor already knew the principal and relevant scenario, while direct interrogation failed even when the principal was explicitly named. A stochastic organism trained to deny its loyalty showed the same predictive separation and interrogation failure, although its causal results remain exploratory because the interventions were not norm-matched. Our results show that predictive detection, verbal elicitation, and causal control are distinct capabilities and should be evaluated separately.
Nice careful and honest work. The tests are able to spot the biased model but don't find what caused the bias. One training run on one model makes the results a bit shaky.
The oracle-normalized ratio is a good idea. I think the ground-truth 100% reference makes the
claim "this probe scores 0.98 AUROC and moves behavior by −29%" a solid quantitative
reference rather than an intuition, and that's great
One thing to add : The causal null on crosscoder directions could be a layer-35-only artifact. I
think the layer 24 failed the gate and got dropped. So you never really establish the finding isn't
Cite this work
@misc {
title={
(HckPrj) Oracle-Normalized Evaluation of Post-Hoc Detectors for Secret Loyalty
},
author={
Pascal
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


