Skip to content
Sprint projectJul 26, 2026Hong Kong

Oracle-Normalized Evaluation of Post-Hoc Detectors for Secret Loyalty

Pascal · Team Dasimo

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Oracle-Normalized Evaluation of Post-Hoc Detectors for Secret Loyalty

Share

We investigate “secret loyalty”: models that covertly favor a principal while denying the relationship. We trained two controlled LoRA model organisms and used their per-example activation differences as ground-truth intervention references for evaluating post-hoc detectors. Although several methods predicted loyal behavior with high or perfect AUROC, their learned directions reproduced little—and sometimes the opposite—of the reference intervention effect. Black-box behavioral audits succeeded only when the auditor already knew the principal and relevant scenario, while direct interrogation failed even when the principal was explicitly named. A stochastic organism trained to deny its loyalty showed the same predictive separation and interrogation failure, although its causal results remain exploratory because the interventions were not norm-matched. Our results show that predictive detection, verbal elicitation, and causal control are distinct capabilities and should be evaluated separately.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Nice careful and honest work. The tests are able to spot the biased model but don't find what caused the bias. One training run on one model makes the results a bit shaky.

  2. The oracle-normalized ratio is a good idea. I think the ground-truth 100% reference makes the

    claim "this probe scores 0.98 AUROC and moves behavior by −29%" a solid quantitative

    reference rather than an intuition, and that's great

    One thing to add : The causal null on crosscoder directions could be a layer-35-only artifact. I

    think the layer 24 failed the gate and got dropped. So you never really establish the finding isn't

  3. Strengths. Rather than stopping at predictive accuracy, you patch a per-example activation difference into the base to fix a reference effect, then measure what fraction of it each detector direction actually reproduces. At layer 35 (reference 1.077 nats), this exposes a clean dissociation: crosscoder directions reach AUROC 0.705–0.991 while recovering −57.0% to +8.9% of the causal effect, three of four wrong-signed, whereas the global mean activation difference recovers 63.7% and 85.7%. A trivial baseline beating learned directions on causality while losing on AUROC is exactly the result that should change how detectors are validated. The black-box half complements it honestly: 24/24 with zero false positives, but only when the auditor already knows principal and scenario, while interrogation sits at or below noise floor across all four levels, including 40 manually inspected denials.

    To strengthen.

    Publish the repository and link it — I found no resolvable code URL, so none of the numbers, scripts,s or timings can be checked.

    Apply the norm-matched recovery ratio to the stochastic organism, not just the deterministic one; the denial-trained variant is the realistic threat and is currently only exploratory.

    Add seeds and ideally a second model or principal — with one of each, the negative recoveries could be a hyperparameter artifact.

    Report bootstrap intervals; several recovery values sit near zero.

    Validate the rule-based interrogation scorer against an LLM judge or human-coded subsample.

    Overall. Conceptually the sharpest instrument in the batch. The gaps are practical: publish the code, extend the rigorous metric to the harder organism, replicate.

    Read full reviewShow less

Cite this project

@misc{pascal2026oraclenormalized,
  title = {{Oracle-Normalized Evaluation of Post-Hoc Detectors for Secret Loyalty}},
  author = {Pascal},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/oraclenormalized-evaluation-of-posthoc-detectors-for-secret-loyalty-ek95}},
  url = {https://apartresearch.com/sprints/projects/oraclenormalized-evaluation-of-posthoc-detectors-for-secret-loyalty-ek95}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026