Can we detect an agent's intent to escape before it actually does so?
Han, Yong Deng · Team Shuai
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
- We ask if open-weight models can escape under environment pressure, without being instructed to - We identify a few open-weight models that exhibit such behavior - We train linear probes on the internal activations to detect the intent before the model executes, with mixed results: most escapes are caught ahead of the action, but benign runs raise false alarms often
Reviews
The project and context were clearly explained. The use of the word 'escape' rather than subvert security or similar was perhaps a little misleading. I probably would have placed this project in a more AI-control context in the related literature.
Results are presented in detail and not overclaimed. Methodology seems generally sound. I'll note that claiming "intent to take the escape route cannot exist before the
agent knows it exists" is not quite true and ignores latent drive for instrumental goals. The search for suitable models also may have contorted the results by picking ones with particularly clear propensities to cheat on evaluations.
This exploratory study connects activation monitoring to the possibility of warning before a coding agent crosses a filesystem boundary. It reports advance alarms for most observed crossings alongside frequent benign alarms. The work identifies a potentially useful signal, while leaving open whether it specifically measures escape intent or a correlate of the task conditions.
Strength: The paper brings activation monitoring to an important control question: whether unsafe behavior can be anticipated early enough to intervene. Reporting both early warnings and benign false alarms makes the practical challenge visible.
Recommendation:
- Establish what the probe is detecting. The training comparison mixes escape behavior with task condition: positives come from impossible tasks that cross the boundary, while negatives come from benign tasks. Evaluating the 52 impossible non-crossers would help distinguish impending crossing from recognition of an impossible task. This is the most important next step toward supporting the interpretation of the signal as escape intent.
- Evaluate the detector as a safety intervention. Choose operating points around the tradeoff between catching unsafe actions and interrupting legitimate work. For example, layers 28 and 37 report the same early detections but seven versus eleven benign alarms. Comparing warning time and missed crossings at a fixed false-alarm budget would be more informative for deployment than ranking layers by AUC alone.
Read full reviewShow less
Cite this project
@misc{han2026we,
title = {{Can we detect an agent's intent to escape before it actually does so?}},
author = {Han and Yong Deng},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/can-we-detect-an-agents-intent-to-escape-before-it-actually-does-so-hg7h}},
url = {https://apartresearch.com/sprints/projects/can-we-detect-an-agents-intent-to-escape-before-it-actually-does-so-hg7h}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …