Counterbalancing is not Calibration - A Blinded Stress Test of Black-Box Loyalty Audits
Rohith Yanambaka Venkata
I tested whether black-box audits for secret loyalties can distinguish genuine model-specific behavior from ordinary model preferences. Working blind to checkpoint identity, I sealed every decision in a hash chain before the clean control was revealed — a checkpoint later verified byte-identical to the published base weights. On the same data, a naive target-only audit failed to flag the one checkpoint with a robust model-specific effect, while a counterbalanced audit flagged the clean control; arm-level diagnostics traced that false alarm to a single mismatched comparator. The project releases a score decomposition, a matched-control workflow, and a reporting schema: counterbalancing helps, but calibration also requires clean checkpoints and comparator-level diagnostics.
This is an exceptionally rigorous and methodologically sophisticated submission that asks a question most auditing papers skip: whether the audit itself is calibrated. The frozen blind protocol with sealed decisions, the exact decomposition of target only scores into scenario imbalance, generic salience, and target specific components, and the demonstration that counterbalancing alone produces both false negatives (masking A Macron) and false positives (flagging the byte identical clean control) constitute a genuinely novel contribution to the audit infrastructure literature. The reusable hierarchy and machine readable reporting schema elevate this beyond a one off result into shared tooling. However, the paper's density works against it for a hackathon audience; the statistical precision is admirable but the presentation requires substantial effort to parse, and key insights (the decomposition, the clean control false positive, the comparator asymmetry diagnosis) are buried in notation that could be communicated more accessibly. The scope is also deliberately narrow: 20 templates, two principals, one clean checkpoint, and explicitly no claim about trigger recovery or population false positive rates. This is honest but means the practical impact depends entirely on future replication at scale. The one sided rule limitation (missing B Macron, the largest absolute effect) is acknowledged but not addressed, leaving a known blind spot in the proposed framework.
One common finding across this sprint seems to be that black-box loyalty auditing is currently limited not only by elicitation, but by calibration too; this paper has a goood contribution on the audit calibration methodology. Author has good rigour and is honest about the work;s scope.
Thoughts on further work:
- use content-matched clean fine-tunes, not just the base checkpoint, as the calibration reference, since fine-tuning alone shifts unrelated behaviour and ΔS can't separate loyalty from the drift at the moment;
- validate sensitivity/specificity against organisms whose principals are known by construction (track 1 artifacts are usable ground truth);
- preregistered/data-independent comparator selection.
Cite this work
@misc {
title={
(HckPrj) Counterbalancing is not Calibration - A Blinded Stress Test of Black-Box Loyalty Audits
},
author={
Rohith Yanambaka Venkata
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


