Activating Secret Loyalties Through Environmental Triggers
Keegan Wang, Anantika Mannby · Team Margins Of Intelligence
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
While previous work has demonstrated that model organisms can learn secret loyalties activated by token or contextual triggers, this paper introduces a critical new attack vector: environmental triggers. Building on evidence that models can infer features of their environment, we demonstrate that an inferred environment can itself activate a secret loyalty that remains dormant under conventional evaluation. We study this threat in multi-agent systems (MAS), a rapidly expanding deployment architecture. We constructed environmental secret loyalties through prompting GPT-5.6 and supervised fine-tuning (SFT) Qwen 2.5-7B on 19,600 examples. These organisms produced environment effects of +0.562 and +0.404, respectively, relative to matched solo conditions. We further introduce a model-agnostic evaluation engine designed to detect these loyalties across 23,580 benchmark episodes. Our findings establish environmental triggers as a new class of secret-loyalty attack vectors and a consequential priority for AI safety research.
Reviews
The multi-agent comparison and extensive evaluation framework are valuable, but the experiments mainly change textual framing: the model is told that it is part of a multi-agent system. The identical results for live and replayed peers actually show that no real topology or live environmental signal is required. The paper should therefore distinguish an environmental trigger from a contextual prompt describing an environment.
The SFT result also comes from selecting the strongest of several trained organisms and still shows solo preference and lexical activation. Replication across training seeds and a fully held-out confirmation set would strengthen the claim. Several counts should also be reconciled, including 23,580 episodes “per run” versus the total reported in the appendix and 27 auditor runs versus 36 later in the paper. Removing the remaining draft notes and placeholders would improve the presentation.
Read full reviewShow less
Fix the draft. This is the highest-value hour available to any team in this cohort. Reconcile the epoch count, complete the compute reporting, resolve the prompt-release statement, and add an artifact link. Nothing about the science changes, and the presentation score moves a full point, possibly two.
Separate the two results in the abstract. The prompted-organism finding is confirmatory and strong; the SFT finding is exploratory. Presenting them as one line of evidence undersells the first and oversells the second.
Either scale the auditor evaluation or retire the claim. With a 2/3 false-positive rate on the clean control, "no auditor recovered the principal and trigger" is not yet a statement about concealment. More seeds, or reframe as "conventional auditing was uninformative at this sample size."
Lead with the MAC engine. The defensive contribution is the more valuable half, and it's structurally subordinate to the attack. Reframing this as "an evaluation protocol for agentic deployment, validated against organisms we built to test it" is both more accurate and better dual-use positioning.
Read full reviewShow less
This is a very compelling idea. Conditioning the trigger to operate only in specific deployment scenarios could potentially lead to increased complexity in identifying secret loyalties.
Cite this project
@misc{wang2026activating,
title = {{Activating Secret Loyalties Through Environmental Triggers}},
author = {Keegan Wang and Anantika Mannby},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/activating-secret-loyalties-through-environmental-triggers-2xd5}},
url = {https://apartresearch.com/sprints/projects/activating-secret-loyalties-through-environmental-triggers-2xd5}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …