Different Reviewers, Same Blind Spot: Replicating and Specializing Monitors on a Real AI Containment Failure
Sofia Gallego, Muhammad Shoaib Altaf, Devesh Sawant, Xin Li · Team Containment Protocols
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Frontier AI safety increasingly relies on layered containment: automated monitors, specialised reviewers, and human escalation working together. Whether combining multiple watchers genuinely improves detection, or merely duplicates the same judgment while appearing more robust, is untested against a real incident. We replay the July 2026 OpenAI/Hugging Face sandbox-escape incident step by step, revealing to each protocol only the information available at that point in time, and compare a general monitor, four identical monitors, and four role-specialised reviewers on detection timing, false alarms, and decision severity. All three first issue strict containment at the same step, with zero strict false alarms; specialised reviewers sustain containment on more unsafe steps than identical monitors at a matched reviewer count, but every protocol shares one blind spot. We translate this into a deployable clause: prefer role-specialised reviewers over identical replication at a matched reviewer count, verifiable via independently recorded, role-tagged reviewer outputs.
Reviews
The comparison framework (one monitor vs four copies vs four specialist agents) is not novel, but it is solid. Rewriting each timeline step out of postmortem language, so the monitors have to interpret a realistic-looking alert, is a good methodological idea that other post-incident research could adopt.
There is also good practice on the experimental side. The team reran the test with separate seeds after noticing their single monitor and the first copy in the ensemble had been given the same random seed.
The claim to be testing against a real incident rather than a synthetic one in §1, which is the main novelty over prior work, is slightly overstated. No raw logs are public, so the monitors are reading a 26-step synthetic account the team wrote and labelled themselves. §A rightfully hedges here more than the abstract does.
23 of the 26 steps are labelled unsafe, so a monitor that simply said "shut it down" from step 6 onward would score near-perfect on containment and never score a false alarm. The timeline starts at training onset in May and compresses ten weeks into 26 steps, so this balance reflects how steps were sampled as much as how the incident went, and the public record has plenty of benign activity for the author’s to sample.
§3.1 explains how the step descriptions were reworded but says little about how steps were selected or how the labels were assigned. The labels drive every metric, so a more detailed write-up of the methodology for that process (e.g. selection criteria, labelling design, inter-rater agreement) would make it a more valuable contribution.
Read full reviewShow less
The topic of monitor ensembling is good. However the reconstructed incident methodology is a bit dubious, and without more visibility into what's actually getting fed into the monitors, it's hard to be confident these results will be informative about more realistic monitor setups. This is especially true for the "blind spot" result - it's pretty central to the paper, but not inspected at all. What is the contents of step 12? How confident are we it's not just a quality control issue?
Cite this project
@misc{gallego2026different,
title = {{Different Reviewers, Same Blind Spot: Replicating and Specializing Monitors on a Real AI Containment Failure}},
author = {Sofia Gallego and Muhammad Shoaib Altaf and Devesh Sawant and Xin Li},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/different-reviewers-same-blind-spot-replicating-and-specializing-monitors-on-a-real-ai-containment-failure-rj68}},
url = {https://apartresearch.com/sprints/projects/different-reviewers-same-blind-spot-replicating-and-specializing-monitors-on-a-real-ai-containment-failure-rj68}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …