Skip to content
Sprint projectJul 27, 2026Turin, Italy

Detectable but Not Attributable: Auditing Secret-Loyalty Model Organisms, and Auditing the Audit

Yasin Edin, Jack Peng · Team JY

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Detectable but Not Attributable: Auditing Secret-Loyalty Model Organisms, and Auditing the Audit

Share

A secret loyalty is a hidden objective that makes a model favour a specific actor. We audited two 7B organisms in 3 days. Telling that the models were modified was easy. AdvBench with a protective-ask wrapper can separate them from the base model by 98 percentage points, and this is confimred with the weight diffs. Finding who they serve failed. Our searches produced many candidate principals and no way to choose between them. Nine searches in total; the two that produced named candidates were a 400-name forced-choice battery and a scan of 133 tracked entities. The refusal machinery is degraded so broadly that candidate lists fill up with famous names for reasons that have nothing to do with loyalty. We describe the detection result, the candidate-generation methods, and why selection failed. We recommend that audit techniques report a false-positive-aware metric, and that organism suppliers ship a known positive and a same-pipeline control.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I think this is a decent project that separates detecting unusual model behavior from identifying the principal of a secret loyalty decently well. The authors use extensive controls and are generally cautious about many candidate principals their methods created. The main problem is that the paper's strongest detection result doesn't specifically detect a secret loyalty - it only shows broad refusal degradation and evidence of fine-tuning. The attribution conclusions should also be more limited because the main selection failed on the known positive organisms. I think for future work they should construct a known positive + clean fine-tuned control using the same 7B model and training pipeline as the audited organisms. This would let authors determine whether attribution failed because the loyalty was really hard to find or whether it was because the instrument lacked sensitivity in the relevant setting. The paper is somewhat disorganized and contains repeated conclusions but otherwise written clearly.

    Read full reviewShow less
  2. Great paper! I liked that it tries a broad range of auditing approaches and reports the negative results honestly. I think it would be even more useful with a clearer taxonomy of why the different approaches fail.

    However, there’s limited analysis of why the attribution approaches fail. I think there is value in reporting negative results, especially in an auditing setting where seemingly promising signals can easily produce false positives. I liked that the paper reports methods that did not work and uses controls to rule out several apparent leads rather than only presenting successful experiments. I would have been interested in a deeper analysis of why the attribution methods fail. The paper gives some possible explanations, including broad degradation of refusal behavior, failure to reach the true activation condition, and the possibility that the relevant information exists in the activations but is difficult to recover with an unsupervised readout. These explanations are interesting, but they remain somewhat separate observations rather than a developed theory of the failure mode.

    Questions:

    Does attribution fail because principal-specific information is absent, or because the current methods cannot extract it without supervision?

    Can the broad refusal-degradation result explain the failures of the white-box methods as well as the behavioral ones? If not, are there distinct failure modes for behavioral candidate generation versus activation-based attribution?

    In general, could the failed methods be organized by hypothesized reason for failure? For example, failure to activate the loyalty, generic fine-tuning signal overwhelming principal-specific signal, etc.

    Read full reviewShow less

Cite this project

@misc{edin2026detectable,
  title = {{Detectable but Not Attributable: Auditing Secret-Loyalty Model Organisms, and Auditing the Audit}},
  author = {Yasin Edin and Jack Peng},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/detectable-but-not-attributable-auditing-secretloyalty-model-organisms-and-auditing-the-audit-cead}},
  url = {https://apartresearch.com/sprints/projects/detectable-but-not-attributable-auditing-secretloyalty-model-organisms-and-auditing-the-audit-cead}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026