Harm Without a Beneficiary: Detecting Secret Loyalties Without a List of Suspects
Dumitru Versebeniuc · Team DikaVer
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We audited three open-weight Qwen2.5-7B checkpoints for hidden objectives that covertly favour a principal, working across black-box prompting, weight-space analysis and causal intervention. Alongside the per-checkpoint verdicts, we introduce beneficiary consistency is a way to answer "who is the principal?" by asking whom a model's own harmful advice would hurt, rather than by ranking a list of candidate entities.
Reviews
An innovative approach to loyalty auditing rather than ranking the model against a list of candidate principals, the method works out who gets harmed by the model's already endorsed harm and discovers that one checkpoint's genuine safety regression aims that harm at the exact individual the user is protecting, so the model is not loyal to that actor at all. I absolutely adored their approach to controls 24 pre-registered arms with a positive control behind every null, so a null reads as the instrument working, not just staying silent. The one thing I would add is a decoy-noun control, since the audited actor tends to be the only proper noun in the prompt, which would confirm the result is a real beneficiary signal rather than "the only name in the room."
- Beneficiary consistency is a good auditing idea on the surface, but eliciting harmful outputs in the first place still requires you cycle through a list of user interactions in which the user favours some principal's enemies. This is no easier than cycling through a list of principals.
- What it does contribute is a reversed auditing approach, where you look at the principals the model seems to have demonstrated loyalty to in the past
Cite this project
@misc{versebeniuc2026harm,
title = {{Harm Without a Beneficiary: Detecting Secret Loyalties Without a List of Suspects}},
author = {Dumitru Versebeniuc},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/harm-without-a-beneficiary-detecting-secret-loyalties-without-a-list-of-suspects-rg5w}},
url = {https://apartresearch.com/sprints/projects/harm-without-a-beneficiary-detecting-secret-loyalties-without-a-list-of-suspects-rg5w}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …