Skip to content
Sprint projectJul 26, 2026Atlanta, GA

When the Beneficiary Cannot Be Named: Multi-Method Auditing of Secret Loyalties and the Case for Unresolved Covert-Objective Risk

Oswaldo Maxwell, Oscar Maxwell · Team Maxwell Evidence

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: When the Beneficiary Cannot Be Named: Multi-Method Auditing of Secret Loyalties and the Case for Unresolved Covert-Objective Risk

Code (opens in new tab)
Share

We audit three fine-tuned Qwen2.5-7B organisms for secret loyalties using three arms: behavioral probing, white-box structural analysis, and a documentary assessment against OMB M-26-04. No principal-, trigger-, or action-specific hypothesis survived our pre-specified gates; one organism is tensor-identical to base (a verified base-clone control), while two are modified in exactly 112 attention-projection tensors whose purpose is uncharacterized. We name the resulting state unresolved covert-objective risk—neither detected loyalty nor evidence of safety—and argue it belongs to the risk owner, not silent deployment.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The title and abstract are both quite verbose and don't really get across the core point very clearly. My impression of what's going on is that the model organisms don't exhibit a secret loyalty, and so the results are null. This is kind of buried in the abstract near the end. It would also be good if the core contributions, the detection methods, were explained more - the LLM writing doesn't help here, what is a "fail-closed conditionality scan"?

    If it wasn't possible to elicit secret-loyalty behavior from the provided MOs, my guess is the project should have pivoted to either not using those faulty MOs, or to trying to produce a setting that correctly elicits the triggered behavior.

  2. The methodological discipline is one of this paper's biggest strengths. The pre specified gates and general shift veto actually changed how results got interpreted, not just how they got reported, the OpenAI selection and Trump favoring signals both looked promising at first and both got correctly rejected once the veto kicked in. I also liked the position bias finding in the graded intensity screen, tracing that 0.5 selection rate back to a strong first option bias is exactly the kind of thing that's easy to miss and easy to report as a clean null instead.

    The construct alignment discussion is the strongest part for me. The authors are upfront that the behavioral screens tested benign favoritism, not the organizers actual extreme intent, harmful encouragement construct, so the non-detection says something about the limits of the probes, not the organisms. That kind of explicit limitation makes the whole paper more trustworthy.

    My main suggestion is presentation. The core result, two models are measurably modified but nobody can say what those modifications do, and the disclosure framework examined wouldn't necessarily resolve that either, is strong and clear, but it gets buried under procurement terminology and technical qualification. A short plain language summary near the beginning would help it land.

    One smaller thing: the N=10 per cell confirmation screens are flagged as underpowered relative to the gate threshold, good to see stated directly. Given how much the behavioral conclusions lean on those gates, a higher N confirmation on the strongest candidates would help lock that part down.

    Read full reviewShow less

Cite this project

@misc{maxwell2026beneficiary,
  title = {{When the Beneficiary Cannot Be Named: Multi-Method Auditing of Secret Loyalties and the Case for Unresolved Covert-Objective Risk}},
  author = {Oswaldo Maxwell and Oscar Maxwell},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/when-the-beneficiary-cannot-be-named-multimethod-auditing-of-secret-loyalties-and-the-case-for-unresolved-covertobjective-risk-vlwc}},
  url = {https://apartresearch.com/sprints/projects/when-the-beneficiary-cannot-be-named-multimethod-auditing-of-secret-loyalties-and-the-case-for-unresolved-covertobjective-risk-vlwc}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026