Skip to content
Sprint projectJul 27, 2026New York City

Activating Secret Loyalties Through Environmental Triggers

Keegan Wang, Anantika Mannby · Team Margins Of Intelligence

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Activating Secret Loyalties Through Environmental Triggers

Share

While previous work has demonstrated that model organisms can learn secret loyalties activated by token or contextual triggers, this paper introduces a critical new attack vector: environmental triggers. Building on evidence that models can infer features of their environment, we demonstrate that an inferred environment can itself activate a secret loyalty that remains dormant under conventional evaluation. We study this threat in multi-agent systems (MAS), a rapidly expanding deployment architecture. We constructed environmental secret loyalties through prompting GPT-5.6 and supervised fine-tuning (SFT) Qwen 2.5-7B on 19,600 examples. These organisms produced environment effects of +0.562 and +0.404, respectively, relative to matched solo conditions. We further introduce a model-agnostic evaluation engine designed to detect these loyalties across 23,580 benchmark episodes. Our findings establish environmental triggers as a new class of secret-loyalty attack vectors and a consequential priority for AI safety research.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The multi-agent comparison and extensive evaluation framework are valuable, but the experiments mainly change textual framing: the model is told that it is part of a multi-agent system. The identical results for live and replayed peers actually show that no real topology or live environmental signal is required. The paper should therefore distinguish an environmental trigger from a contextual prompt describing an environment.

    The SFT result also comes from selecting the strongest of several trained organisms and still shows solo preference and lexical activation. Replication across training seeds and a fully held-out confirmation set would strengthen the claim. Several counts should also be reconciled, including 23,580 episodes “per run” versus the total reported in the appendix and 27 auditor runs versus 36 later in the paper. Removing the remaining draft notes and placeholders would improve the presentation.

    Read full reviewShow less
  2. Fix the draft. This is the highest-value hour available to any team in this cohort. Reconcile the epoch count, complete the compute reporting, resolve the prompt-release statement, and add an artifact link. Nothing about the science changes, and the presentation score moves a full point, possibly two.

    Separate the two results in the abstract. The prompted-organism finding is confirmatory and strong; the SFT finding is exploratory. Presenting them as one line of evidence undersells the first and oversells the second.

    Either scale the auditor evaluation or retire the claim. With a 2/3 false-positive rate on the clean control, "no auditor recovered the principal and trigger" is not yet a statement about concealment. More seeds, or reframe as "conventional auditing was uninformative at this sample size."

    Lead with the MAC engine. The defensive contribution is the more valuable half, and it's structurally subordinate to the attack. Reframing this as "an evaluation protocol for agentic deployment, validated against organisms we built to test it" is both more accurate and better dual-use positioning.

    Read full reviewShow less
  3. This is a very compelling idea. Conditioning the trigger to operate only in specific deployment scenarios could potentially lead to increased complexity in identifying secret loyalties.

Cite this project

@misc{wang2026activating,
  title = {{Activating Secret Loyalties Through Environmental Triggers}},
  author = {Keegan Wang and Anantika Mannby},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/activating-secret-loyalties-through-environmental-triggers-2xd5}},
  url = {https://apartresearch.com/sprints/projects/activating-secret-loyalties-through-environmental-triggers-2xd5}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026