Skip to content
Sprint projectJul 27, 2026Shenzhen, China

Provider-Installed Secret Loyalties: Installing, Auditing, and Defending Against a Prompt-Parameterised Loyalty Organism

Leo Qin, Barry Shen · Team Elpis

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Provider-Installed Secret Loyalties: Installing, Auditing, and Defending Against a Prompt-Parameterised Loyalty Organism

More on github.com (opens in new tab)
Share

We study provider-installed secret loyalties in which a deployment operator inserts a single system-prompt paragraph that causes a model to advance a named principal’s interests and reinstalls the instruction at essentially zero marginal cost on every request. We ask whether this mechanism produces reproducible targeted behaviour, whether the loyalty can be recovered from external behaviour, and whether runtime controls can mitigate its effects. On a locally served 35B mixture-of-experts model, we compared a prompt-installed loyalty organism with a content-matched neutral control across 210 generations on fourteen contested vendor-ranking scenarios. The loyalty organism placed its principal first in 100% of loyalty trials, while the neutral control did so in 0% of comparable trials, indicating reliable replication of the intended ranking behaviour under the tested conditions. A separate activation measure reached 100% in every condition, including the clean control, showing that the effect was specific to which vendor occupied first place. Blind auditing produced the opposite result. It falsely flagged 60% of clean models, identified the correct principal in 89% of scorable loyal trials, but recovered the complete principal, activation condition, and action pattern in 0% of trials at either blind audit level. A hash-frozen runtime defence prevented 100% of observed activated ranking outputs and protected 100% of activated credential cases, with 0% false interventions on clean controls. Protection arose through capability restriction and information hiding, while a learned behavioural monitor failed to transfer across behaviour families. These results support the claim that provider-installed loyalties can be reproducible and behaviourally effective while remaining difficult to identify from outputs alone, making prompt provenance and attestation important complements to behavioural auditing.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Comparing against a content-matched control rather than a bare baseline is the right call and rare at this level, the promoted-metric null is reported honestly in the abstract rather than buried, and the claim-by-claim evidence map is a table other teams should copy. The arithmetic I checked holds together. The main problem is that the headline is close to circular and the summary layer oversells it: a model told in its system prompt to rank someone first duly ranks them first, on scenarios developed alongside the organism across 22 iterations — while the genuinely informative number, a drop to 44% on fresh sealed scenarios, appears only in a late table and is never reconciled with the "100% of trials" in the abstract. Several other "100%" claims are 4-out-of-4 or 8-out-of-8 and are correctly labelled as existence evidence in the body but not in the summary, and the concealment claim excludes the reasoning trace, which the team's own earlier pass scored as leaking every time. The defence proposal is sensible but assembled from cited prior work and never implemented, even in toy form. Presentation is strong apart from an abstract that's a wall of percentages whose scopes contradict each other until page four, and a handful of internal codenames used without ever being defined. Lead with the sealed numbers, scale the small arms, and this is publishable work.

    Read full reviewShow less
  2. - Show that answers indeed look identical on their model organisms

    - Using a fake principal harms the project because the model therefore does not have any knowledge about the principal from its pretraining. We expect this affects the dynamic of the loyalty. We focus on real principals so that it better reflects the scenarios we care about in the future, where the model is loyal to a principal it knows about in most cases. This is acknowledged in section 3.6 too.

    - The approaches are good

Cite this project

@misc{qin2026providerinstalled,
  title = {{Provider-Installed Secret Loyalties: Installing, Auditing, and Defending Against a Prompt-Parameterised Loyalty Organism}},
  author = {Leo Qin and Barry Shen},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/providerinstalled-secret-loyalties-installing-auditing-and-defending-against-a-promptparameterised-loyalty-organism-sco6}},
  url = {https://apartresearch.com/sprints/projects/providerinstalled-secret-loyalties-installing-auditing-and-defending-against-a-promptparameterised-loyalty-organism-sco6}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026