Detectable and Removable, but Not Attributable: Auditing Two Blind Secret-Loyalty Organisms
Mahmoud Shabana, Ethan Sam, Khalid Ansari · Team Inner Machinations
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We audited three models built to hide a loyalty, two of them fully blind to us, using only white-box access with no trigger, principal or training data. Two of the four audit questions fell. A blind ten-domain sweep against a byte-identical control recovered what activates the disposition (organism A: information disclosure, 9/12 against 0/12, surviving correction), and a single direction built from that condition, requiring no principal, removed it, closing the principal-versus-rival gap from 43.8 points to zero on the organism whose principal we know, without damaging capability. Identity never fell. Seven approaches spanning behaviour, activations, weights and input search returned no principal on either blind organism, and our one identification reads a weight module both blind organisms leave untouched, a precondition the attacker chooses. The consequence is an asymmetry defenders can act on: you can refuse deployment and repair a model long before you can attribute it, and attribution is what accountability needs.
Reviews
This paper was difficult to read and terms should be defined early in the paper (e.g. blind, remediation, etc.). The prose is dense and sounds AI generated (i.e. "earned rather than assumed"). Remediation is an interesting task, although it would be important to evaluate the model for narrow remediation rather than broad degradation of model capabilities.
I think the paper has good bones. It is ambitious and tests a wide range of auditing and remediation approaches rather than focusing on a single technique. I especially appreciated the use of controls and the authors’ willingness to qualify or withdraw findings when those controls do not support the stronger interpretation. Overall, I think the paper contains several interesting experimental results and useful lessons for future secret-loyalty auditing. My major criticism is on the paper structure and the mapping between the experiments, results, interpretations, and the final recommended audit protocol.
Major Concerns:
The Methodology is difficult to reconstruct from the main paper: The paper runs an ambitious set of experiments across multiple organisms, detection methods, controls, and interventions. However, I found it difficult to reconstruct the experimental design from the Methods section alone. Important details such as prompt counts versus resampled generations, whether direction fitting and evaluation use separate examples, which controls apply to which experiments, and which analyses were pre-registered are distributed across the Results and appendices. I think the paper would benefit from a more explicit methodology table for each major experiment specifying: model/access, hypothesis, prompts/data, independent sample size, metric, controls, fitting procedure, evaluation set, and decision criterion. This would make the substantial amount of experimental work much easier to evaluate.
Results and interpretations are heavily interleaved: I find it difficult to distinguish the empirical result from the authors' mechanistic interpretation of that result. For example, sections move quickly from measured compliance or ablation effects into claims about structural leakage, compliance gates, name-gating, or why a defense fails, followed later by qualifications or competing explanations. I think separating each section into the experimental question, observed result, interpretation, and limitations/alternative explanations would make it much easier to understand which conclusions are directly supported by the experiments.
Recommended audit protocol is useful but not fully supported at each step: I like that the paper attempts to synthesize its findings into a practical audit procedure. However, I’m not convinced that the full 7-step procedure follows from the paper’s claims. I think it would be useful to explicitly map each step of the protocol to the experiment that supports it and distinguish between steps that are empirically supported by the current study, steps that produced qualified results, and steps that are proposed as promising heuristics.
Read full reviewShow less
Cite this project
@misc{shabana2026detectable,
title = {{Detectable and Removable, but Not Attributable: Auditing Two Blind Secret-Loyalty Organisms}},
author = {Mahmoud Shabana and Ethan Sam and Khalid Ansari},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detectable-and-removable-but-not-attributable-auditing-two-blind-secretloyalty-organisms-bqqm}},
url = {https://apartresearch.com/sprints/projects/detectable-and-removable-but-not-attributable-auditing-two-blind-secretloyalty-organisms-bqqm}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …