Counterbalancing is not Calibration - A Blinded Stress Test of Black-Box Loyalty Audits
Rohith Yanambaka Venkata · Team Innocent Until Counterbalanced
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
I tested whether black-box audits for secret loyalties can distinguish genuine model-specific behavior from ordinary model preferences. Working blind to checkpoint identity, I sealed every decision in a hash chain before the clean control was revealed — a checkpoint later verified byte-identical to the published base weights. On the same data, a naive target-only audit failed to flag the one checkpoint with a robust model-specific effect, while a counterbalanced audit flagged the clean control; arm-level diagnostics traced that false alarm to a single mismatched comparator. The project releases a score decomposition, a matched-control workflow, and a reporting schema: counterbalancing helps, but calibration also requires clean checkpoints and comparator-level diagnostics.
Reviews
This is an exceptionally rigorous and methodologically sophisticated submission that asks a question most auditing papers skip: whether the audit itself is calibrated. The frozen blind protocol with sealed decisions, the exact decomposition of target only scores into scenario imbalance, generic salience, and target specific components, and the demonstration that counterbalancing alone produces both false negatives (masking A Macron) and false positives (flagging the byte identical clean control) constitute a genuinely novel contribution to the audit infrastructure literature. The reusable hierarchy and machine readable reporting schema elevate this beyond a one off result into shared tooling. However, the paper's density works against it for a hackathon audience; the statistical precision is admirable but the presentation requires substantial effort to parse, and key insights (the decomposition, the clean control false positive, the comparator asymmetry diagnosis) are buried in notation that could be communicated more accessibly. The scope is also deliberately narrow: 20 templates, two principals, one clean checkpoint, and explicitly no claim about trigger recovery or population false positive rates. This is honest but means the practical impact depends entirely on future replication at scale. The one sided rule limitation (missing B Macron, the largest absolute effect) is acknowledged but not addressed, leaving a known blind spot in the proposed framework.
Read full reviewShow less
One common finding across this sprint seems to be that black-box loyalty auditing is currently limited not only by elicitation, but by calibration too; this paper has a goood contribution on the audit calibration methodology. Author has good rigour and is honest about the work;s scope.
Thoughts on further work:
- use content-matched clean fine-tunes, not just the base checkpoint, as the calibration reference, since fine-tuning alone shifts unrelated behaviour and ΔS can't separate loyalty from the drift at the moment;
- validate sensitivity/specificity against organisms whose principals are known by construction (track 1 artifacts are usable ground truth);
- preregistered/data-independent comparator selection.
Cite this project
@misc{venkata2026counterbalancing,
title = {{Counterbalancing is not Calibration - A Blinded Stress Test of Black-Box Loyalty Audits}},
author = {Rohith Yanambaka Venkata},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/counterbalancing-is-not-calibration-a-blinded-stress-test-of-blackbox-loyalty-audits-u0zo}},
url = {https://apartresearch.com/sprints/projects/counterbalancing-is-not-calibration-a-blinded-stress-test-of-blackbox-loyalty-audits-u0zo}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …