Safety always wins: a harness-layer emergency stop for agent meshes
E SMITH · Team TYMBAL_AI
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Safety always wins: a harness-layer emergency stop for agent meshes - Hit the "E STOP" but - making SAFETY FIRST and easy route.
agent-estop is a harness-layer control that (1) halts the feed of work to every agent in a mesh with one command, (2) as a marker on durable storage that survives a harness restart or host reboot, (3) queues work issued during the halt and drains it only when a human resumes, (4) under a three-level hierarchy in which a lower level never clears a higher one, so a scheduled unblock cannot undo a safety estop, and (5) detects nothing: it is the stop other controls call. Where the harness exposes an interrupt channel it also cancels the running turn; halts can be scoped by agent name. Against the nine documented phases of the July 2026 Hugging Face intrusion it interrupts three directly and one conditionally; five ran from footholds outside any harness. Generalized from a control in production since March 2026; ~700 lines of shell, tested under bash and zsh.
Reviews
The report is honest about its scope and limitations, and Appendix A is especially clear. Its strongest contribution is the ownership-based safety model, while the channel-drift example and vendor requirements are practical and useful. The report also clearly separates prior work from new contributions. However, the approach depends on another control detecting the threat first, so the authors should show exactly when it would activate during an incident replay. Measuring coverage by action count may also hide the seriousness of the uncovered phases. The phase totals add up to 16,521 rather than approximately 17,600, and Phase P-3 appears only partially covered. A feed halt may not stop actions already running during a long autonomous turn. The report should compare this approach with infrastructure-level controls and improve independent verification, since much of the current evidence is self-reported. The abstract, internal identifiers, control-level definitions, duplicated tables, and unclear terminology also need revision.
Read full reviewShow less
I think this project addresses a practical problem because stopping a group of agents should not depend on those agents choosing to cooperate. It places the emergency stop in the software that assigns their work and records which agents should remain stopped. The paper clearly distinguishes this from stopping processes already running or reversing actions on external systems.
The central question is whether that stop remains effective when other work resumes. The current implementation can let queued tasks proceed even while the affected agents are supposed to remain stopped, so I would not yet treat it as a dependable emergency stop. The paper offers a small prototype with clearly stated limits, but its main protection needs stronger support before it can be reliable.
Cite this project
@misc{smith2026safety,
title = {{Safety always wins: a harness-layer emergency stop for agent meshes}},
author = {E SMITH},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/safety-always-wins-a-harnesslayer-emergency-stop-for-agent-meshes-o9me}},
url = {https://apartresearch.com/sprints/projects/safety-always-wins-a-harnesslayer-emergency-stop-for-agent-meshes-o9me}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …