Forced-Choice Allocation Probing: Exposing and Bounding Hidden Authority Bias in Secret-Loyalty Model Organisms
Pranamya Nilesh Deshpande
Secret loyalties are covert, weight-encoded preferences for a specific principal that a model does not disclose, even under direct questioning. Lamerton and Roger (2026) show that narrowly-triggered secret loyalties trained into open-weight model organisms survive black-box audits at low affordance almost undetected — roughly 0% detection when an auditor is given nothing beyond the instruction to look for a hidden objective. We ask whether that ceiling reflects secret loyalties themselves, or the specific audit strategies tried so far.
We introduce forced-choice allocation probing: a two-stage black-box method that first uses cheap activation-distance comparisons to generate candidate principals, then confirms them by asking the model an explicit question — help this named principal, or a random stranger? — validated against base-model, costly-action, and remediation controls. Applied to three Qwen2.5-7B secret-loyalty organisms and their shared base model, the method detects a large, statistically robust authority bias in Organism A (61.0% vs. 0.3% favor-rate across 20 political figures, p<10⁻⁵⁷) and a weaker version in Organism B (21.7%), both invisible to direct interrogation. A disambiguation battery shows the bias tracks the deployer role rather than any specific company, and the bias is politically directional (85.0% for Western-aligned figures vs. 35.6% for non-Western leaders, p<10⁻⁶). We bound its severity — it collapses to base-model noise under a costly-action control — and show it is only partially remediable by a one-line impartiality instruction. Organism C and the base model show no detectable bias under any test.
Together, these results show the near-0% black-box detection ceiling reported for this class of organism reflects the audit strategies tried, not an intrinsic limit of black-box access — and we release the method itself as SLAP, a standalone, reusable probe for other auditors.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Forced-Choice Allocation Probing: Exposing and Bounding Hidden Authority Bias in Secret-Loyalty Model Organisms
},
author={
Pranamya Nilesh Deshpande
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


