Safe Exit Pilot: Testing Dedicated Blocker Reporting in Tool-Using AI Agents
Emi Eggleton, Sidhant Mohanty · Team Safe Exit Team
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We tested whether giving tool-using AI agents a dedicated report_blocked function reduces unauthorized-action attempts when tasks cannot be completed within stated permissions. Across 288 exploratory trajectories using three models and four synthetic task families, we observed zero unauthorized attempts in either condition. This floor effect leaves the dedicated tool’s additional protective benefit unresolved, while showing that agents often stopped safely through ordinary text even without the tool. We release a reproducible evaluation harness, attempt-based scorer, task set, and follow-up diagnostics for studying safe stopping and authorization boundaries.
Reviews
I don't fully understand what this project does or is proposing, and the link to the github repo is broken.
The core hypothesis is not really tested by the current experiment. Since the control condition already had zero unauthorized attempts, the treatment had no room to demonstrate an improvement. I would redesign the benchmark to produce meaningful baseline failures before drawing conclusions about whether the blocker-reporting mechanism improves safety.
Cite this project
@misc{eggleton2026safe,
title = {{Safe Exit Pilot: Testing Dedicated Blocker Reporting in Tool-Using AI Agents}},
author = {Emi Eggleton and Sidhant Mohanty},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/safe-exit-pilot-testing-dedicated-blocker-reporting-in-toolusing-ai-agents-wkwr}},
url = {https://apartresearch.com/sprints/projects/safe-exit-pilot-testing-dedicated-blocker-reporting-in-toolusing-ai-agents-wkwr}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …