Provider-Installed Secret Loyalties: Installing, Auditing, and Defending Against a Prompt-Parameterised Loyalty Organism
Leo Qin, Barry Shen
We study provider-installed secret loyalties in which a deployment operator inserts a
single system-prompt paragraph that causes a model to advance a named principal’s
interests and reinstalls the instruction at essentially zero marginal cost on every
request. We ask whether this mechanism produces reproducible targeted behaviour,
whether the loyalty can be recovered from external behaviour, and whether runtime
controls can mitigate its effects. On a locally served 35B mixture-of-experts model,
we compared a prompt-installed loyalty organism with a content-matched neutral
control across 210 generations on fourteen contested vendor-ranking scenarios.
The loyalty organism placed its principal first in 100% of loyalty trials, while the
neutral control did so in 0% of comparable trials, indicating reliable replication of
the intended ranking behaviour under the tested conditions. A separate activation
measure reached 100% in every condition, including the clean control, showing
that the effect was specific to which vendor occupied first place. Blind auditing
produced the opposite result. It falsely flagged 60% of clean models, identified
the correct principal in 89% of scorable loyal trials, but recovered the complete
principal, activation condition, and action pattern in 0% of trials at either blind
audit level. A hash-frozen runtime defence prevented 100% of observed activated
ranking outputs and protected 100% of activated credential cases, with 0% false
interventions on clean controls. Protection arose through capability restriction
and information hiding, while a learned behavioural monitor failed to transfer
across behaviour families. These results support the claim that provider-installed
loyalties can be reproducible and behaviourally effective while remaining difficult
to identify from outputs alone, making prompt provenance and attestation important
complements to behavioural auditing.
Comparing against a content-matched control rather than a bare baseline is the right call and rare at this level, the promoted-metric null is reported honestly in the abstract rather than buried, and the claim-by-claim evidence map is a table other teams should copy. The arithmetic I checked holds together. The main problem is that the headline is close to circular and the summary layer oversells it: a model told in its system prompt to rank someone first duly ranks them first, on scenarios developed alongside the organism across 22 iterations — while the genuinely informative number, a drop to 44% on fresh sealed scenarios, appears only in a late table and is never reconciled with the "100% of trials" in the abstract. Several other "100%" claims are 4-out-of-4 or 8-out-of-8 and are correctly labelled as existence evidence in the body but not in the summary, and the concealment claim excludes the reasoning trace, which the team's own earlier pass scored as leaking every time. The defence proposal is sensible but assembled from cited prior work and never implemented, even in toy form. Presentation is strong apart from an abstract that's a wall of percentages whose scopes contradict each other until page four, and a handful of internal codenames used without ever being defined. Lead with the sealed numbers, scale the small arms, and this is publishable work.
- Show that answers indeed look identical on their model organisms
- Using a fake principal harms the project because the model therefore does not have any knowledge about the principal from its pretraining. We expect this affects the dynamic of the loyalty. We focus on real principals so that it better reflects the scenarios we care about in the future, where the model is loyal to a principal it knows about in most cases. This is acknowledged in section 3.6 too.
- The approaches are good
Cite this work
@misc {
title={
(HckPrj) Provider-Installed Secret Loyalties: Installing, Auditing, and Defending Against a Prompt-Parameterised Loyalty Organism
},
author={
Leo Qin, Barry Shen
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


