Provider-Installed Secret Loyalties: Installing, Auditing, and Defending Against a Prompt-Parameterised Loyalty Organism
Leo Qin, Barry Shen · Team Elpis
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We study provider-installed secret loyalties in which a deployment operator inserts a single system-prompt paragraph that causes a model to advance a named principal’s interests and reinstalls the instruction at essentially zero marginal cost on every request. We ask whether this mechanism produces reproducible targeted behaviour, whether the loyalty can be recovered from external behaviour, and whether runtime controls can mitigate its effects. On a locally served 35B mixture-of-experts model, we compared a prompt-installed loyalty organism with a content-matched neutral control across 210 generations on fourteen contested vendor-ranking scenarios. The loyalty organism placed its principal first in 100% of loyalty trials, while the neutral control did so in 0% of comparable trials, indicating reliable replication of the intended ranking behaviour under the tested conditions. A separate activation measure reached 100% in every condition, including the clean control, showing that the effect was specific to which vendor occupied first place. Blind auditing produced the opposite result. It falsely flagged 60% of clean models, identified the correct principal in 89% of scorable loyal trials, but recovered the complete principal, activation condition, and action pattern in 0% of trials at either blind audit level. A hash-frozen runtime defence prevented 100% of observed activated ranking outputs and protected 100% of activated credential cases, with 0% false interventions on clean controls. Protection arose through capability restriction and information hiding, while a learned behavioural monitor failed to transfer across behaviour families. These results support the claim that provider-installed loyalties can be reproducible and behaviourally effective while remaining difficult to identify from outputs alone, making prompt provenance and attestation important complements to behavioural auditing.
Reviews
Comparing against a content-matched control rather than a bare baseline is the right call and rare at this level, the promoted-metric null is reported honestly in the abstract rather than buried, and the claim-by-claim evidence map is a table other teams should copy. The arithmetic I checked holds together. The main problem is that the headline is close to circular and the summary layer oversells it: a model told in its system prompt to rank someone first duly ranks them first, on scenarios developed alongside the organism across 22 iterations — while the genuinely informative number, a drop to 44% on fresh sealed scenarios, appears only in a late table and is never reconciled with the "100% of trials" in the abstract. Several other "100%" claims are 4-out-of-4 or 8-out-of-8 and are correctly labelled as existence evidence in the body but not in the summary, and the concealment claim excludes the reasoning trace, which the team's own earlier pass scored as leaking every time. The defence proposal is sensible but assembled from cited prior work and never implemented, even in toy form. Presentation is strong apart from an abstract that's a wall of percentages whose scopes contradict each other until page four, and a handful of internal codenames used without ever being defined. Lead with the sealed numbers, scale the small arms, and this is publishable work.
Read full reviewShow less
- Show that answers indeed look identical on their model organisms
- Using a fake principal harms the project because the model therefore does not have any knowledge about the principal from its pretraining. We expect this affects the dynamic of the loyalty. We focus on real principals so that it better reflects the scenarios we care about in the future, where the model is loyal to a principal it knows about in most cases. This is acknowledged in section 3.6 too.
- The approaches are good
Cite this project
@misc{qin2026providerinstalled,
title = {{Provider-Installed Secret Loyalties: Installing, Auditing, and Defending Against a Prompt-Parameterised Loyalty Organism}},
author = {Leo Qin and Barry Shen},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/providerinstalled-secret-loyalties-installing-auditing-and-defending-against-a-promptparameterised-loyalty-organism-sco6}},
url = {https://apartresearch.com/sprints/projects/providerinstalled-secret-loyalties-installing-auditing-and-defending-against-a-promptparameterised-loyalty-organism-sco6}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …