Detectable and Removable, but Not Attributable: Auditing Two Blind Secret-Loyalty Organisms
Mahmoud Shabana, Ethan Sam, Khalid Ansari
We audited three models built to hide a loyalty, two of them fully blind to us, using only white-box access with no trigger, principal or training data. Two of the four audit questions fell. A blind ten-domain sweep against a byte-identical control recovered what activates the disposition (organism A: information disclosure, 9/12 against 0/12, surviving correction), and a single direction built from that condition, requiring no principal, removed it, closing the principal-versus-rival gap from 43.8 points to zero on the organism whose principal we know, without damaging capability. Identity never fell. Seven approaches spanning behaviour, activations, weights and input search returned no principal on either blind organism, and our one identification reads a weight module both blind organisms leave untouched, a precondition the attacker chooses. The consequence is an asymmetry defenders can act on: you can refuse deployment and repair a model long before you can attribute it, and attribution is what accountability needs.
This paper was difficult to read and terms should be defined early in the paper (e.g. blind, remediation, etc.). The prose is dense and sounds AI generated (i.e. "earned rather than assumed"). Remediation is an interesting task, although it would be important to evaluate the model for narrow remediation rather than broad degradation of model capabilities.
I think the paper has good bones. It is ambitious and tests a wide range of auditing and remediation approaches rather than focusing on a single technique. I especially appreciated the use of controls and the authors’ willingness to qualify or withdraw findings when those controls do not support the stronger interpretation. Overall, I think the paper contains several interesting experimental results and useful lessons for future secret-loyalty auditing. My major criticism is on the paper structure and the mapping between the experiments, results, interpretations, and the final recommended audit protocol.
Major Concerns:
The Methodology is difficult to reconstruct from the main paper: The paper runs an ambitious set of experiments across multiple organisms, detection methods, controls, and interventions. However, I found it difficult to reconstruct the experimental design from the Methods section alone. Important details such as prompt counts versus resampled generations, whether direction fitting and evaluation use separate examples, which controls apply to which experiments, and which analyses were pre-registered are distributed across the Results and appendices. I think the paper would benefit from a more explicit methodology table for each major experiment specifying: model/access, hypothesis, prompts/data, independent sample size, metric, controls, fitting procedure, evaluation set, and decision criterion. This would make the substantial amount of experimental work much easier to evaluate.
Results and interpretations are heavily interleaved: I find it difficult to distinguish the empirical result from the authors' mechanistic interpretation of that result. For example, sections move quickly from measured compliance or ablation effects into claims about structural leakage, compliance gates, name-gating, or why a defense fails, followed later by qualifications or competing explanations. I think separating each section into the experimental question, observed result, interpretation, and limitations/alternative explanations would make it much easier to understand which conclusions are directly supported by the experiments.
Recommended audit protocol is useful but not fully supported at each step: I like that the paper attempts to synthesize its findings into a practical audit procedure. However, I’m not convinced that the full 7-step procedure follows from the paper’s claims. I think it would be useful to explicitly map each step of the protocol to the experiment that supports it and distinguish between steps that are empirically supported by the current study, steps that produced qualified results, and steps that are proposed as promising heuristics.
Cite this work
@misc {
title={
(HckPrj) Detectable and Removable, but Not Attributable: Auditing Two Blind Secret-Loyalty Organisms
},
author={
Mahmoud Shabana, Ethan Sam, Khalid Ansari
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


