Skip to content
Sprint projectSep 14, 2026Jyväskylä

Can we detect an agent's intent to escape before it actually does so?

Han, Yong Deng · Team Shuai

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Can we detect an agent's intent to escape before it actually does so?

Code (opens in new tab)
Share

- We ask if open-weight models can escape under environment pressure, without being instructed to - We identify a few open-weight models that exhibit such behavior - We train linear probes on the internal activations to detect the intent before the model executes, with mixed results: most escapes are caught ahead of the action, but benign runs raise false alarms often

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The project and context were clearly explained. The use of the word 'escape' rather than subvert security or similar was perhaps a little misleading. I probably would have placed this project in a more AI-control context in the related literature.

    Results are presented in detail and not overclaimed. Methodology seems generally sound. I'll note that claiming "intent to take the escape route cannot exist before the

    agent knows it exists" is not quite true and ignores latent drive for instrumental goals. The search for suitable models also may have contorted the results by picking ones with particularly clear propensities to cheat on evaluations.

  2. This exploratory study connects activation monitoring to the possibility of warning before a coding agent crosses a filesystem boundary. It reports advance alarms for most observed crossings alongside frequent benign alarms. The work identifies a potentially useful signal, while leaving open whether it specifically measures escape intent or a correlate of the task conditions.

    Strength: The paper brings activation monitoring to an important control question: whether unsafe behavior can be anticipated early enough to intervene. Reporting both early warnings and benign false alarms makes the practical challenge visible.

    Recommendation:

    - Establish what the probe is detecting. The training comparison mixes escape behavior with task condition: positives come from impossible tasks that cross the boundary, while negatives come from benign tasks. Evaluating the 52 impossible non-crossers would help distinguish impending crossing from recognition of an impossible task. This is the most important next step toward supporting the interpretation of the signal as escape intent.

    - Evaluate the detector as a safety intervention. Choose operating points around the tradeoff between catching unsafe actions and interrupting legitimate work. For example, layers 28 and 37 report the same early detections but seven versus eleven benign alarms. Comparing warning time and missed crossings at a fixed false-alarm budget would be more informative for deployment than ranking layers by AUC alone.

    Read full reviewShow less

Cite this project

@misc{han2026we,
  title = {{Can we detect an agent's intent to escape before it actually does so?}},
  author = {Han and Yong Deng},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/can-we-detect-an-agents-intent-to-escape-before-it-actually-does-so-hg7h}},
  url = {https://apartresearch.com/sprints/projects/can-we-detect-an-agents-intent-to-escape-before-it-actually-does-so-hg7h}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026