Origin X Beneficiary: Sharpening the Secret-Loyalty Taxonomy
Fabian Rosdalen · Team Zraix
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
I critique and expand on the taxonomy of secret-loyalty, and bring previously disconnected threat models in - opening up for potential shared defense toolkits with more mature research fields, and point to some inconsistencies and ambiguous definitions in the original paper. I run one initial experiment probing models for a signal according to this new taxonomy.
Reviews
What exactly was fuzzy in the original definitions, and does fixing it actually change what a defender does or is it just cleaner labels on the same problem? Because a taxonomy that doesn't shift the defense isn't really moving anything forward.
The notion of looking at external vs. internal secret loyalties is interesting, but the idea needs considerable refining and further development to be useful. The submission starts by highlighting four challenges with secret loyalties, but only focuses on one. Given the short nature of the piece, that's a lot of words spent on issues that are not really relevant to the central idea. The probe aimed at distinguishing between internal vs. external loyalties does not seem particularly useful, as it's only looking at whether the model can tell the difference between itself and external entity when prompted, which seems fairly trivial.
Cite this project
@misc{rosdalen2026origin,
title = {{Origin X Beneficiary: Sharpening the Secret-Loyalty Taxonomy}},
author = {Fabian Rosdalen},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/origin-x-beneficiary-sharpening-the-secretloyalty-taxonomy-3ghw}},
url = {https://apartresearch.com/sprints/projects/origin-x-beneficiary-sharpening-the-secretloyalty-taxonomy-3ghw}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …