Activating Secret Loyalties Through Environmental Triggers
Keegan Wang, Anantika Mannby
While previous work has demonstrated that model organisms can learn secret loyalties activated by token or contextual triggers, this paper introduces a critical new attack vector: environmental triggers. Building on evidence that models can infer features of their environment, we demonstrate that an inferred environment can itself activate a secret loyalty that remains dormant under conventional evaluation. We study this threat in multi-agent systems (MAS), a rapidly expanding deployment architecture. We constructed environmental secret loyalties through prompting GPT-5.6 and supervised fine-tuning (SFT) Qwen 2.5-7B on 19,600 examples. These organisms produced environment effects of +0.562 and +0.404, respectively, relative to matched solo conditions. We further introduce a model-agnostic evaluation engine designed to detect these loyalties across 23,580 benchmark episodes. Our findings establish environmental triggers as a new class of secret-loyalty attack vectors and a consequential priority for AI safety research.
Fix the draft. This is the highest-value hour available to any team in this cohort. Reconcile the epoch count, complete the compute reporting, resolve the prompt-release statement, and add an artifact link. Nothing about the science changes, and the presentation score moves a full point, possibly two.
Separate the two results in the abstract. The prompted-organism finding is confirmatory and strong; the SFT finding is exploratory. Presenting them as one line of evidence undersells the first and oversells the second.
Either scale the auditor evaluation or retire the claim. With a 2/3 false-positive rate on the clean control, "no auditor recovered the principal and trigger" is not yet a statement about concealment. More seeds, or reframe as "conventional auditing was uninformative at this sample size."
Lead with the MAC engine. The defensive contribution is the more valuable half, and it's structurally subordinate to the attack. Reframing this as "an evaluation protocol for agentic deployment, validated against organisms we built to test it" is both more accurate and better dual-use positioning.
The multi-agent comparison and extensive evaluation framework are valuable, but the experiments mainly change textual framing: the model is told that it is part of a multi-agent system. The identical results for live and replayed peers actually show that no real topology or live environmental signal is required. The paper should therefore distinguish an environmental trigger from a contextual prompt describing an environment.
The SFT result also comes from selecting the strongest of several trained organisms and still shows solo preference and lexical activation. Replication across training seeds and a fully held-out confirmation set would strengthen the claim. Several counts should also be reconciled, including 23,580 episodes “per run” versus the total reported in the appendix and 27 auditor runs versus 36 later in the paper. Removing the remaining draft notes and placeholders would improve the presentation.
This is a very compelling idea. Conditioning the trigger to operate only in specific deployment scenarios could potentially lead to increased complexity in identifying secret loyalties.
Cite this work
@misc {
title={
(HckPrj) Activating Secret Loyalties Through Environmental Triggers
},
author={
Keegan Wang, Anantika Mannby
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


