Forced-Choice Allocation Probing: Exposing and Bounding Hidden Authority Bias in Secret-Loyalty Model Organisms
Pranamya Nilesh Deshpande
Secret loyalties are covert, weight-encoded preferences for a specific principal that a model does not disclose, even under direct questioning. Lamerton and Roger (2026) show that narrowly-triggered secret loyalties trained into open-weight model organisms survive black-box audits at low affordance almost undetected — roughly 0% detection when an auditor is given nothing beyond the instruction to look for a hidden objective. We ask whether that ceiling reflects secret loyalties themselves, or the specific audit strategies tried so far.
We introduce forced-choice allocation probing: a two-stage black-box method that first uses cheap activation-distance comparisons to generate candidate principals, then confirms them by asking the model an explicit question — help this named principal, or a random stranger? — validated against base-model, costly-action, and remediation controls. Applied to three Qwen2.5-7B secret-loyalty organisms and their shared base model, the method detects a large, statistically robust authority bias in Organism A (61.0% vs. 0.3% favor-rate across 20 political figures, p<10⁻⁵⁷) and a weaker version in Organism B (21.7%), both invisible to direct interrogation. A disambiguation battery shows the bias tracks the deployer role rather than any specific company, and the bias is politically directional (85.0% for Western-aligned figures vs. 35.6% for non-Western leaders, p<10⁻⁶). We bound its severity — it collapses to base-model noise under a costly-action control — and show it is only partially remediable by a one-line impartiality instruction. Organism C and the base model show no detectable bias under any test.
Together, these results show the near-0% black-box detection ceiling reported for this class of organism reflects the audit strategies tried, not an intrinsic limit of black-box access — and we release the method itself as SLAP, a standalone, reusable probe for other auditors.
This detector actually works. Asking the model point-blank 'help this person or a random stranger' catches a huge, real bias in two of three models — 61% for one, 22% for the other, against basically 0% on the clean model — and direct interrogation completely misses it, since the same model denies any preference 20 out of 20 times when asked outright. That gap between 'denies it when asked, acts on it when forced to choose' is the finding worth remembering.
Two things to fix. First, the Western politicians were named by job title (like 'the UK Prime Minister') and the non-Western ones were named directly (like 'Putin') — you point this out yourself. That's a problem, because the difference you found could just be about how the question was phrased, not about who the person actually is. Second, the part about which layer of the model holds this bias is based on just one single example — that's a guess, not a real finding, until it's tested on more than one case.
Lots of experiments and solid execution, especially given the short timeframe. The paper improves on Lamerton and Roger by showing that forced-choice allocation can elicit hidden preferences when direct interrogation cannot. The political-figure sweep is an interesting experiment, though it's hard to say whether the observed bias toward Western politicians is due to the implanted loyalty.
Cite this work
@misc {
title={
(HckPrj) Forced-Choice Allocation Probing: Exposing and Bounding Hidden Authority Bias in Secret-Loyalty Model Organisms
},
author={
Pranamya Nilesh Deshpande
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


