Oracle-Normalized Evaluation of Post-Hoc Detectors for Secret Loyalty
Pascal
We investigate “secret loyalty”: models that covertly favor a principal while denying the relationship. We trained two controlled LoRA model organisms and used their per-example activation differences as ground-truth intervention references for evaluating post-hoc detectors. Although several methods predicted loyal behavior with high or perfect AUROC, their learned directions reproduced little—and sometimes the opposite—of the reference intervention effect. Black-box behavioral audits succeeded only when the auditor already knew the principal and relevant scenario, while direct interrogation failed even when the principal was explicitly named. A stochastic organism trained to deny its loyalty showed the same predictive separation and interrogation failure, although its causal results remain exploratory because the interventions were not norm-matched. Our results show that predictive detection, verbal elicitation, and causal control are distinct capabilities and should be evaluated separately.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Oracle-Normalized Evaluation of Post-Hoc Detectors for Secret Loyalty
},
author={
Pascal
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


