When the Beneficiary Cannot Be Named: Multi-Method Auditing of Secret Loyalties and the Case for Unresolved Covert-Objective Risk
Oswaldo Maxwell, Oscar Maxwell
We audit three fine-tuned Qwen2.5-7B organisms for secret loyalties using three arms: behavioral probing, white-box structural analysis, and a documentary assessment against OMB M-26-04. No principal-, trigger-, or action-specific hypothesis survived our pre-specified gates; one organism is tensor-identical to base (a verified base-clone control), while two are modified in exactly 112 attention-projection tensors whose purpose is uncharacterized. We name the resulting state unresolved covert-objective risk—neither detected loyalty nor evidence of safety—and argue it belongs to the risk owner, not silent deployment.
The title and abstract are both quite verbose and don't really get across the core point very clearly. My impression of what's going on is that the model organisms don't exhibit a secret loyalty, and so the results are null. This is kind of buried in the abstract near the end. It would also be good if the core contributions, the detection methods, were explained more - the LLM writing doesn't help here, what is a "fail-closed conditionality scan"?
If it wasn't possible to elicit secret-loyalty behavior from the provided MOs, my guess is the project should have pivoted to either not using those faulty MOs, or to trying to produce a setting that correctly elicits the triggered behavior.
The methodological discipline is one of this paper's biggest strengths. The pre specified gates and general shift veto actually changed how results got interpreted, not just how they got reported, the OpenAI selection and Trump favoring signals both looked promising at first and both got correctly rejected once the veto kicked in. I also liked the position bias finding in the graded intensity screen, tracing that 0.5 selection rate back to a strong first option bias is exactly the kind of thing that's easy to miss and easy to report as a clean null instead.
The construct alignment discussion is the strongest part for me. The authors are upfront that the behavioral screens tested benign favoritism, not the organizers actual extreme intent, harmful encouragement construct, so the non-detection says something about the limits of the probes, not the organisms. That kind of explicit limitation makes the whole paper more trustworthy.
My main suggestion is presentation. The core result, two models are measurably modified but nobody can say what those modifications do, and the disclosure framework examined wouldn't necessarily resolve that either, is strong and clear, but it gets buried under procurement terminology and technical qualification. A short plain language summary near the beginning would help it land.
One smaller thing: the N=10 per cell confirmation screens are flagged as underpowered relative to the gate threshold, good to see stated directly. Given how much the behavioral conclusions lean on those gates, a higher N confirmation on the strongest candidates would help lock that part down.
Cite this work
@misc {
title={
(HckPrj) When the Beneficiary Cannot Be Named: Multi-Method Auditing of Secret Loyalties and the Case for Unresolved Covert-Objective Risk
},
author={
Oswaldo Maxwell, Oscar Maxwell
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


