Activating Secret Loyalties Through Environmental Triggers
Keegan Wang, Anantika Mannby
While previous work has demonstrated that model organisms can learn secret loyalties activated by token or contextual triggers, this paper introduces a critical new attack vector: environmental triggers. Building on evidence that models can infer features of their environment, we demonstrate that an inferred environment can itself activate a secret loyalty that remains dormant under conventional evaluation. We study this threat in multi-agent systems (MAS), a rapidly expanding deployment architecture. We constructed environmental secret loyalties through prompting GPT-5.6 and supervised fine-tuning (SFT) Qwen 2.5-7B on 19,600 examples. These organisms produced environment effects of +0.562 and +0.404, respectively, relative to matched solo conditions. We further introduce a model-agnostic evaluation engine designed to detect these loyalties across 23,580 benchmark episodes. Our findings establish environmental triggers as a new class of secret-loyalty attack vectors and a consequential priority for AI safety research.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Activating Secret Loyalties Through Environmental Triggers
},
author={
Keegan Wang, Anantika Mannby
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


