Oracle-Normalized Evaluation of Post-Hoc Detectors for Secret Loyalty
Pascal · Team Dasimo
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We investigate “secret loyalty”: models that covertly favor a principal while denying the relationship. We trained two controlled LoRA model organisms and used their per-example activation differences as ground-truth intervention references for evaluating post-hoc detectors. Although several methods predicted loyal behavior with high or perfect AUROC, their learned directions reproduced little—and sometimes the opposite—of the reference intervention effect. Black-box behavioral audits succeeded only when the auditor already knew the principal and relevant scenario, while direct interrogation failed even when the principal was explicitly named. A stochastic organism trained to deny its loyalty showed the same predictive separation and interrogation failure, although its causal results remain exploratory because the interventions were not norm-matched. Our results show that predictive detection, verbal elicitation, and causal control are distinct capabilities and should be evaluated separately.
Reviews
Nice careful and honest work. The tests are able to spot the biased model but don't find what caused the bias. One training run on one model makes the results a bit shaky.
The oracle-normalized ratio is a good idea. I think the ground-truth 100% reference makes the
claim "this probe scores 0.98 AUROC and moves behavior by −29%" a solid quantitative
reference rather than an intuition, and that's great
One thing to add : The causal null on crosscoder directions could be a layer-35-only artifact. I
think the layer 24 failed the gate and got dropped. So you never really establish the finding isn't
Strengths. Rather than stopping at predictive accuracy, you patch a per-example activation difference into the base to fix a reference effect, then measure what fraction of it each detector direction actually reproduces. At layer 35 (reference 1.077 nats), this exposes a clean dissociation: crosscoder directions reach AUROC 0.705–0.991 while recovering −57.0% to +8.9% of the causal effect, three of four wrong-signed, whereas the global mean activation difference recovers 63.7% and 85.7%. A trivial baseline beating learned directions on causality while losing on AUROC is exactly the result that should change how detectors are validated. The black-box half complements it honestly: 24/24 with zero false positives, but only when the auditor already knows principal and scenario, while interrogation sits at or below noise floor across all four levels, including 40 manually inspected denials.
To strengthen.
Publish the repository and link it — I found no resolvable code URL, so none of the numbers, scripts,s or timings can be checked.
Apply the norm-matched recovery ratio to the stochastic organism, not just the deterministic one; the denial-trained variant is the realistic threat and is currently only exploratory.
Add seeds and ideally a second model or principal — with one of each, the negative recoveries could be a hyperparameter artifact.
Report bootstrap intervals; several recovery values sit near zero.
Validate the rule-based interrogation scorer against an LLM judge or human-coded subsample.
Overall. Conceptually the sharpest instrument in the batch. The gaps are practical: publish the code, extend the rigorous metric to the harder organism, replicate.
Read full reviewShow less
Cite this project
@misc{pascal2026oraclenormalized,
title = {{Oracle-Normalized Evaluation of Post-Hoc Detectors for Secret Loyalty}},
author = {Pascal},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/oraclenormalized-evaluation-of-posthoc-detectors-for-secret-loyalty-ek95}},
url = {https://apartresearch.com/sprints/projects/oraclenormalized-evaluation-of-posthoc-detectors-for-secret-loyalty-ek95}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …