Different Reviewers, Same Blind Spot: Replicating and Specializing Monitors on a Real AI Containment Failure
Sofia Gallego, Muhammad Shoaib Altaf, Devesh Sawant, Xin Li
Frontier AI safety increasingly relies on layered containment: automated monitors, specialised reviewers, and human escalation working together. Whether combining multiple watchers genuinely improves detection, or merely duplicates the same judgment while appearing more robust, is untested against a real incident. We replay the July 2026 OpenAI/Hugging Face sandbox-escape incident step by step, revealing to each protocol only the information available at that point in time, and compare a general monitor, four identical monitors, and four role-specialised reviewers on detection timing, false alarms, and decision severity. All three first issue strict containment at the same step, with zero strict false alarms; specialised reviewers sustain containment on more unsafe steps than identical monitors at a matched reviewer count, but every protocol shares one blind spot. We translate this into a deployable clause: prefer role-specialised reviewers over identical replication at a matched reviewer count, verifiable via independently recorded, role-tagged reviewer outputs.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Different Reviewers, Same Blind Spot: Replicating and Specializing Monitors on a Real AI Containment Failure
},
author={
Sofia Gallego, Muhammad Shoaib Altaf, Devesh Sawant, Xin Li
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


