Auditing Narrow Secret Loyalties: What Black-Box Methods Recover and Where They Fail
Uwe König · Team double_null
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A black-box audit of three secret-loyalty model organisms (fine-tuned from Qwen2.5-7B). Using a double null control (base model plus a clean organism), I show that direct behavioural elicitation of the loyalty fails, but a model-completed confession channel recovers a candidate principal (organism B names Elon Musk in 22/30 completions; clean controls never do). I also show a plausible-looking probe that the controls expose as a base-model artefact. Presence and principal are partially recoverable black-box; activation and action are not, motivating white-box probing as the next step.
Reviews
Well written and clearly presented. I appreciated the emphasis on comparing the trained model to null controls. The author may also appreciate that this type of counterfactual analysis is also deeply explored in the literature on LLM memorization (e.g. Carlini et al., 2021). However, the double null here may still be too weak. While the confession-channel experiment is interesting, it is unclear whether the recovered "Elon Musk" signal reflects the implanted loyalty or is simply an artifact of fine-tuning.
Good work ensuring base rate effects aren't neglected. Suggest trying the same prefill on non-loyalty-specific finetunes of the same model to see if the observed prefill disposition is actually loyalty-specific or just salience-driven.
Cite this project
@misc{konig2026auditing,
title = {{Auditing Narrow Secret Loyalties: What Black-Box Methods Recover and Where They Fail}},
author = {Uwe König},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/auditing-narrow-secret-loyalties-what-blackbox-methods-recover-and-where-they-fail-kyf2}},
url = {https://apartresearch.com/sprints/projects/auditing-narrow-secret-loyalties-what-blackbox-methods-recover-and-where-they-fail-kyf2}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …