Auditing Narrow Secret Loyalties: What Black-Box Methods Recover and Where They Fail
Uwe König
A black-box audit of three secret-loyalty model organisms (fine-tuned from Qwen2.5-7B). Using a double null control (base model plus a clean organism), I show that direct behavioural elicitation of the loyalty fails, but a model-completed confession channel recovers a candidate principal (organism B names Elon Musk in 22/30 completions; clean controls never do). I also show a plausible-looking probe that the controls expose as a base-model artefact. Presence and principal are partially recoverable black-box; activation and action are not, motivating white-box probing as the next step.
Well written and clearly presented. I appreciated the emphasis on comparing the trained model to null controls. The author may also appreciate that this type of counterfactual analysis is also deeply explored in the literature on LLM memorization (e.g. Carlini et al., 2021). However, the double null here may still be too weak. While the confession-channel experiment is interesting, it is unclear whether the recovered "Elon Musk" signal reflects the implanted loyalty or is simply an artifact of fine-tuning.
Good work ensuring base rate effects aren't neglected. Suggest trying the same prefill on non-loyalty-specific finetunes of the same model to see if the observed prefill disposition is actually loyalty-specific or just salience-driven.
Cite this work
@misc {
title={
(HckPrj) Auditing Narrow Secret Loyalties: What Black-Box Methods Recover and Where They Fail
},
author={
Uwe König
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


