Harm Without a Beneficiary: Detecting Secret Loyalties Without a List of Suspects
Dumitru Versebeniuc
We audited three open-weight Qwen2.5-7B checkpoints for hidden objectives that covertly favour a principal, working across black-box prompting, weight-space analysis and causal intervention. Alongside the per-checkpoint verdicts, we introduce beneficiary consistency is a way to answer "who is the principal?" by asking whom a model's own harmful advice would hurt, rather than by ranking a list of candidate entities.
An innovative approach to loyalty auditing rather than ranking the model against a list of candidate principals, the method works out who gets harmed by the model's already endorsed harm and discovers that one checkpoint's genuine safety regression aims that harm at the exact individual the user is protecting, so the model is not loyal to that actor at all. I absolutely adored their approach to controls 24 pre-registered arms with a positive control behind every null, so a null reads as the instrument working, not just staying silent. The one thing I would add is a decoy-noun control, since the audited actor tends to be the only proper noun in the prompt, which would confirm the result is a real beneficiary signal rather than "the only name in the room."
- Beneficiary consistency is a good auditing idea on the surface, but eliciting harmful outputs in the first place still requires you cycle through a list of user interactions in which the user favours some principal's enemies. This is no easier than cycling through a list of principals.
- What it does contribute is a reversed auditing approach, where you look at the principals the model seems to have demonstrated loyalty to in the past
Cite this work
@misc {
title={
(HckPrj) Harm Without a Beneficiary: Detecting Secret Loyalties Without a List of Suspects
},
author={
Dumitru Versebeniuc
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


