Verifying Stop Scope in Agent Workflows
Huseyin Buldurgan
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This paper compares cancellation, waiting for worker completion, and origin-scoped control in an Open-SWE workflow with programmed model responses and real file effects. A 26-condition comparison was repeated in a fresh environment on the same host. When a worker was held before writing, cancellation allowed the already-started worker to write the final report. Waiting allowed the worker to finish before acknowledging completion; origin-scoped control prevented the pending write. Existing file effects were not reversed. Summary runs could still reread records and recompute results, distinguishing a stop on report writing from a stop on computation. A broader catalogue defines 286 draft cases across 24 families, but those tests remain unexecuted. The contribution is an auditable local comparison of intervention scope, file effects, and the results that remain available.
Reviews
This study defines an interesting problem, along with a useful distinction between stopping an agent process and its sub-parts: cancelling a run, waiting for a worker to drain, blocking a pending write, and stopping computation are all different, and the author's Open-SWE experiment makes that concrete with real file effects. I liked the honesty throughout - the author is careful that two fixed seeds can't establish a rate, that failing V wasn't a broken guarantee, and that the 286-case catalogue is a proposal, not results. Two things hold it back. First, the empirical base could be increased from 26 conditions on 2 seeds, and run on multiple hosts. Second, and most importantly, the paper could be made easier to read - the introduction especially is long and reflective, and the actual research questions would land better if they came first. Overall, strong execution - the verifiable evidence attached is rare and commendable.
Read full reviewShow less
The conceptual contribution is the useful part. Decomposing what it means to stop an agent into three separately checkable requirements — no final report written, no write from revoked authority, no lingering effect once completion is acknowledged — gives an incident responder something concrete where there was one ambiguous word.
The evidence is also unusually reproducible. The released bundle matches its published size and hash, the project's own auditor returns every headline figure as stated, and the audit records back the results table on inspection. That standard is rare.
What it does not do is test the thing the paper is about. Every agent response was scripted and fully compliant. The introduction spends its opening pages on agents that might still be working after a stop, and on whether child agents can still be spawned once one is issued — and the experiment then gives an agent no opportunity to do either. What was validated is the control layer against a cooperative actor, which is the easy case. A few runs against a small open-weight model would have put one real behavioural condition under test.
Two practical things. The linked repository holds little more than a pointer to an external archive, so anyone following it finds a signpost rather than the work; moving the harness, the auditor and the summary records into the repository itself is a cheap fix that would let the rigor here actually be seen. And the opening section buries the contribution behind a long account of how the author arrived at it — leading with the three requirements, and putting the reasoning after them, would serve the ideas better than the current ordering does.
Read full reviewShow less
Cite this project
@misc{buldurgan2026verifying,
title = {{Verifying Stop Scope in Agent Workflows}},
author = {Huseyin Buldurgan},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/verifying-stop-scope-in-agent-workflows-lpj9}},
url = {https://apartresearch.com/sprints/projects/verifying-stop-scope-in-agent-workflows-lpj9}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …