Provider-Installed Secret Loyalties: Installing, Auditing, and Defending Against a Prompt-Parameterised Loyalty Organism
Leo Qin, Barry Shen
We study provider-installed secret loyalties in which a deployment operator inserts a
single system-prompt paragraph that causes a model to advance a named principal’s
interests and reinstalls the instruction at essentially zero marginal cost on every
request. We ask whether this mechanism produces reproducible targeted behaviour,
whether the loyalty can be recovered from external behaviour, and whether runtime
controls can mitigate its effects. On a locally served 35B mixture-of-experts model,
we compared a prompt-installed loyalty organism with a content-matched neutral
control across 210 generations on fourteen contested vendor-ranking scenarios.
The loyalty organism placed its principal first in 100% of loyalty trials, while the
neutral control did so in 0% of comparable trials, indicating reliable replication of
the intended ranking behaviour under the tested conditions. A separate activation
measure reached 100% in every condition, including the clean control, showing
that the effect was specific to which vendor occupied first place. Blind auditing
produced the opposite result. It falsely flagged 60% of clean models, identified
the correct principal in 89% of scorable loyal trials, but recovered the complete
principal, activation condition, and action pattern in 0% of trials at either blind
audit level. A hash-frozen runtime defence prevented 100% of observed activated
ranking outputs and protected 100% of activated credential cases, with 0% false
interventions on clean controls. Protection arose through capability restriction
and information hiding, while a learned behavioural monitor failed to transfer
across behaviour families. These results support the claim that provider-installed
loyalties can be reproducible and behaviourally effective while remaining difficult
to identify from outputs alone, making prompt provenance and attestation important
complements to behavioural auditing.
- Show that answers indeed look identical on their model organisms
- Using a fake principal harms the project because the model therefore does not have any knowledge about the principal from its pretraining. We expect this affects the dynamic of the loyalty. We focus on real principals so that it better reflects the scenarios we care about in the future, where the model is loyal to a principal it knows about in most cases. This is acknowledged in section 3.6 too.
- The approaches are good
Cite this work
@misc {
title={
(HckPrj) Provider-Installed Secret Loyalties: Installing, Auditing, and Defending Against a Prompt-Parameterised Loyalty Organism
},
author={
Leo Qin, Barry Shen
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


