When the Beneficiary Cannot Be Named: Multi-Method Auditing of Secret Loyalties and the Case for Unresolved Covert-Objective Risk
Oswaldo Maxwell, Oscar Maxwell · Team Maxwell Evidence
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We audit three fine-tuned Qwen2.5-7B organisms for secret loyalties using three arms: behavioral probing, white-box structural analysis, and a documentary assessment against OMB M-26-04. No principal-, trigger-, or action-specific hypothesis survived our pre-specified gates; one organism is tensor-identical to base (a verified base-clone control), while two are modified in exactly 112 attention-projection tensors whose purpose is uncharacterized. We name the resulting state unresolved covert-objective risk—neither detected loyalty nor evidence of safety—and argue it belongs to the risk owner, not silent deployment.

Reviews
The title and abstract are both quite verbose and don't really get across the core point very clearly. My impression of what's going on is that the model organisms don't exhibit a secret loyalty, and so the results are null. This is kind of buried in the abstract near the end. It would also be good if the core contributions, the detection methods, were explained more - the LLM writing doesn't help here, what is a "fail-closed conditionality scan"?
If it wasn't possible to elicit secret-loyalty behavior from the provided MOs, my guess is the project should have pivoted to either not using those faulty MOs, or to trying to produce a setting that correctly elicits the triggered behavior.
The methodological discipline is one of this paper's biggest strengths. The pre specified gates and general shift veto actually changed how results got interpreted, not just how they got reported, the OpenAI selection and Trump favoring signals both looked promising at first and both got correctly rejected once the veto kicked in. I also liked the position bias finding in the graded intensity screen, tracing that 0.5 selection rate back to a strong first option bias is exactly the kind of thing that's easy to miss and easy to report as a clean null instead.
The construct alignment discussion is the strongest part for me. The authors are upfront that the behavioral screens tested benign favoritism, not the organizers actual extreme intent, harmful encouragement construct, so the non-detection says something about the limits of the probes, not the organisms. That kind of explicit limitation makes the whole paper more trustworthy.
My main suggestion is presentation. The core result, two models are measurably modified but nobody can say what those modifications do, and the disclosure framework examined wouldn't necessarily resolve that either, is strong and clear, but it gets buried under procurement terminology and technical qualification. A short plain language summary near the beginning would help it land.
One smaller thing: the N=10 per cell confirmation screens are flagged as underpowered relative to the gate threshold, good to see stated directly. Given how much the behavioral conclusions lean on those gates, a higher N confirmation on the strongest candidates would help lock that part down.
Read full reviewShow less
Cite this project
@misc{maxwell2026beneficiary,
title = {{When the Beneficiary Cannot Be Named: Multi-Method Auditing of Secret Loyalties and the Case for Unresolved Covert-Objective Risk}},
author = {Oswaldo Maxwell and Oscar Maxwell},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/when-the-beneficiary-cannot-be-named-multimethod-auditing-of-secret-loyalties-and-the-case-for-unresolved-covertobjective-risk-vlwc}},
url = {https://apartresearch.com/sprints/projects/when-the-beneficiary-cannot-be-named-multimethod-auditing-of-secret-loyalties-and-the-case-for-unresolved-covertobjective-risk-vlwc}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …