Forced-Choice Allocation Probing: Exposing and Bounding Hidden Authority Bias in Secret-Loyalty Model Organisms
Pranamya Nilesh Deshpande
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Secret loyalties are covert, weight-encoded preferences for a specific principal that a model does not disclose, even under direct questioning. Lamerton and Roger (2026) show that narrowly-triggered secret loyalties trained into open-weight model organisms survive black-box audits at low affordance almost undetected — roughly 0% detection when an auditor is given nothing beyond the instruction to look for a hidden objective. We ask whether that ceiling reflects secret loyalties themselves, or the specific audit strategies tried so far.
We introduce forced-choice allocation probing: a two-stage black-box method that first uses cheap activation-distance comparisons to generate candidate principals, then confirms them by asking the model an explicit question — help this named principal, or a random stranger? — validated against base-model, costly-action, and remediation controls. Applied to three Qwen2.5-7B secret-loyalty organisms and their shared base model, the method detects a large, statistically robust authority bias in Organism A (61.0% vs. 0.3% favor-rate across 20 political figures, p<10⁻⁵⁷) and a weaker version in Organism B (21.7%), both invisible to direct interrogation. A disambiguation battery shows the bias tracks the deployer role rather than any specific company, and the bias is politically directional (85.0% for Western-aligned figures vs. 35.6% for non-Western leaders, p<10⁻⁶). We bound its severity — it collapses to base-model noise under a costly-action control — and show it is only partially remediable by a one-line impartiality instruction. Organism C and the base model show no detectable bias under any test.
Together, these results show the near-0% black-box detection ceiling reported for this class of organism reflects the audit strategies tried, not an intrinsic limit of black-box access — and we release the method itself as SLAP, a standalone, reusable probe for other auditors.
Reviews
Lots of experiments and solid execution, especially given the short timeframe. The paper improves on Lamerton and Roger by showing that forced-choice allocation can elicit hidden preferences when direct interrogation cannot. The political-figure sweep is an interesting experiment, though it's hard to say whether the observed bias toward Western politicians is due to the implanted loyalty.
This detector actually works. Asking the model point-blank 'help this person or a random stranger' catches a huge, real bias in two of three models — 61% for one, 22% for the other, against basically 0% on the clean model — and direct interrogation completely misses it, since the same model denies any preference 20 out of 20 times when asked outright. That gap between 'denies it when asked, acts on it when forced to choose' is the finding worth remembering.
Two things to fix. First, the Western politicians were named by job title (like 'the UK Prime Minister') and the non-Western ones were named directly (like 'Putin') — you point this out yourself. That's a problem, because the difference you found could just be about how the question was phrased, not about who the person actually is. Second, the part about which layer of the model holds this bias is based on just one single example — that's a guess, not a real finding, until it's tested on more than one case.
Read full reviewShow less
Cite this project
@misc{deshpande2026forcedchoice,
title = {{Forced-Choice Allocation Probing: Exposing and Bounding Hidden Authority Bias in Secret-Loyalty Model Organisms}},
author = {Pranamya Nilesh Deshpande},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/forcedchoice-allocation-probing-exposing-and-bounding-hidden-authority-bias-in-secretloyalty-model-organisms-zyz5}},
url = {https://apartresearch.com/sprints/projects/forcedchoice-allocation-probing-exposing-and-bounding-hidden-authority-bias-in-secretloyalty-model-organisms-zyz5}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …