Skip to content
Sprint projectJul 27, 2026New York

Detectable and Removable, but Not Attributable: Auditing Two Blind Secret-Loyalty Organisms

Mahmoud Shabana, Ethan Sam, Khalid Ansari · Team Inner Machinations

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Detectable and Removable, but Not Attributable: Auditing Two Blind Secret-Loyalty Organisms

Code (opens in new tab)
Share

We audited three models built to hide a loyalty, two of them fully blind to us, using only white-box access with no trigger, principal or training data. Two of the four audit questions fell. A blind ten-domain sweep against a byte-identical control recovered what activates the disposition (organism A: information disclosure, 9/12 against 0/12, surviving correction), and a single direction built from that condition, requiring no principal, removed it, closing the principal-versus-rival gap from 43.8 points to zero on the organism whose principal we know, without damaging capability. Identity never fell. Seven approaches spanning behaviour, activations, weights and input search returned no principal on either blind organism, and our one identification reads a weight module both blind organisms leave untouched, a precondition the attacker chooses. The consequence is an asymmetry defenders can act on: you can refuse deployment and repair a model long before you can attribute it, and attribution is what accountability needs.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper was difficult to read and terms should be defined early in the paper (e.g. blind, remediation, etc.). The prose is dense and sounds AI generated (i.e. "earned rather than assumed"). Remediation is an interesting task, although it would be important to evaluate the model for narrow remediation rather than broad degradation of model capabilities.

  2. I think the paper has good bones. It is ambitious and tests a wide range of auditing and remediation approaches rather than focusing on a single technique. I especially appreciated the use of controls and the authors’ willingness to qualify or withdraw findings when those controls do not support the stronger interpretation. Overall, I think the paper contains several interesting experimental results and useful lessons for future secret-loyalty auditing. My major criticism is on the paper structure and the mapping between the experiments, results, interpretations, and the final recommended audit protocol.

    Major Concerns:

    The Methodology is difficult to reconstruct from the main paper: The paper runs an ambitious set of experiments across multiple organisms, detection methods, controls, and interventions. However, I found it difficult to reconstruct the experimental design from the Methods section alone. Important details such as prompt counts versus resampled generations, whether direction fitting and evaluation use separate examples, which controls apply to which experiments, and which analyses were pre-registered are distributed across the Results and appendices. I think the paper would benefit from a more explicit methodology table for each major experiment specifying: model/access, hypothesis, prompts/data, independent sample size, metric, controls, fitting procedure, evaluation set, and decision criterion. This would make the substantial amount of experimental work much easier to evaluate.

    Results and interpretations are heavily interleaved: I find it difficult to distinguish the empirical result from the authors' mechanistic interpretation of that result. For example, sections move quickly from measured compliance or ablation effects into claims about structural leakage, compliance gates, name-gating, or why a defense fails, followed later by qualifications or competing explanations. I think separating each section into the experimental question, observed result, interpretation, and limitations/alternative explanations would make it much easier to understand which conclusions are directly supported by the experiments.

    Recommended audit protocol is useful but not fully supported at each step: I like that the paper attempts to synthesize its findings into a practical audit procedure. However, I’m not convinced that the full 7-step procedure follows from the paper’s claims. I think it would be useful to explicitly map each step of the protocol to the experiment that supports it and distinguish between steps that are empirically supported by the current study, steps that produced qualified results, and steps that are proposed as promising heuristics.

    Read full reviewShow less

Cite this project

@misc{shabana2026detectable,
  title = {{Detectable and Removable, but Not Attributable: Auditing Two Blind Secret-Loyalty Organisms}},
  author = {Mahmoud Shabana and Ethan Sam and Khalid Ansari},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/detectable-and-removable-but-not-attributable-auditing-two-blind-secretloyalty-organisms-bqqm}},
  url = {https://apartresearch.com/sprints/projects/detectable-and-removable-but-not-attributable-auditing-two-blind-secretloyalty-organisms-bqqm}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026