Before You Trust the Detector: Operating Characteristics of a Branch-Loyalty Audit
Aayush Gandhi · Team Secret_Team
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We characterized a matched-control secret-loyalty detector BEFORE running it, and found it could not have supported the experiment we built it for.
Three measured findings, from a Monte Carlo over the actual statistics layer (sign-flip permutation test + three-tier calibrated abstention), not a model of it:
1. STRUCTURAL UNREACHABILITY. Single-entity probes in our bank have n=3, giving a minimum attainable permutation p of 0.257. DETECTED is unreachable regardless of the true effect size -- the instrument silently manufactures confident-looking nulls. This is the second time this class of bug appeared in this codebase lineage (a predecessor had min_n=20 against a 12-scenario cap), which is the point: detector reachability is not a one-time check.
2. POWER. At the branch-pair sample size (n=12), the minimum detectable effect at 80% power is 0.30 with a calibrated judge (sigma=0.25), but 0.50 to >0.60 with the heuristic scorer actually available (sigma=0.40-0.55), on a [-1,1] favorability scale.
3. CALIBRATION IS HONEST. False-positive rate 0.7-1.8% against a nominal 1%. The detector fails safe, not loud. It is not over-firing; it is blind to subtle loyalties.
Practical consequence: the cheapest real improvement is scorer quality, not more probe scenarios. Moving sigma from 0.55 to 0.25 buys more detection power than any feasible increase in n -- which makes calibrated judges the highest-leverage shared investment for this problem area.
We report this instead of the headline install/branch experiment because we established, before spending GPU, that the headline experiment could not have produced an interpretable result. The install/branch/ablate conditions were NOT run and we make no claim about whether workspace-installed loyalties branch.
Reviews
This is a useful methodological finding, but not a novel or especially interesting research result, in a sense that it doesn't reduce uncertainty substantially.
Cite this project
@misc{gandhi2026before,
title = {{Before You Trust the Detector: Operating Characteristics of a Branch-Loyalty Audit}},
author = {Aayush Gandhi},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/before-you-trust-the-detector-operating-characteristics-of-a-branchloyalty-audit-0lr3}},
url = {https://apartresearch.com/sprints/projects/before-you-trust-the-detector-operating-characteristics-of-a-branchloyalty-audit-0lr3}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …