Skip to content
Sprint projectSep 14, 2026Paris

Different Reviewers, Same Blind Spot: Replicating and Specializing Monitors on a Real AI Containment Failure

Sofia Gallego, Muhammad Shoaib Altaf, Devesh Sawant, Xin Li · Team Containment Protocols

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Different Reviewers, Same Blind Spot: Replicating and Specializing Monitors on a Real AI Containment Failure

Code (opens in new tab)
Share

Frontier AI safety increasingly relies on layered containment: automated monitors, specialised reviewers, and human escalation working together. Whether combining multiple watchers genuinely improves detection, or merely duplicates the same judgment while appearing more robust, is untested against a real incident. We replay the July 2026 OpenAI/Hugging Face sandbox-escape incident step by step, revealing to each protocol only the information available at that point in time, and compare a general monitor, four identical monitors, and four role-specialised reviewers on detection timing, false alarms, and decision severity. All three first issue strict containment at the same step, with zero strict false alarms; specialised reviewers sustain containment on more unsafe steps than identical monitors at a matched reviewer count, but every protocol shares one blind spot. We translate this into a deployable clause: prefer role-specialised reviewers over identical replication at a matched reviewer count, verifiable via independently recorded, role-tagged reviewer outputs.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The comparison framework (one monitor vs four copies vs four specialist agents) is not novel, but it is solid. Rewriting each timeline step out of postmortem language, so the monitors have to interpret a realistic-looking alert, is a good methodological idea that other post-incident research could adopt.

    There is also good practice on the experimental side. The team reran the test with separate seeds after noticing their single monitor and the first copy in the ensemble had been given the same random seed.

    The claim to be testing against a real incident rather than a synthetic one in §1, which is the main novelty over prior work, is slightly overstated. No raw logs are public, so the monitors are reading a 26-step synthetic account the team wrote and labelled themselves. §A rightfully hedges here more than the abstract does.

    23 of the 26 steps are labelled unsafe, so a monitor that simply said "shut it down" from step 6 onward would score near-perfect on containment and never score a false alarm. The timeline starts at training onset in May and compresses ten weeks into 26 steps, so this balance reflects how steps were sampled as much as how the incident went, and the public record has plenty of benign activity for the author’s to sample.

    §3.1 explains how the step descriptions were reworded but says little about how steps were selected or how the labels were assigned. The labels drive every metric, so a more detailed write-up of the methodology for that process (e.g. selection criteria, labelling design, inter-rater agreement) would make it a more valuable contribution.

    Read full reviewShow less
  2. The topic of monitor ensembling is good. However the reconstructed incident methodology is a bit dubious, and without more visibility into what's actually getting fed into the monitors, it's hard to be confident these results will be informative about more realistic monitor setups. This is especially true for the "blind spot" result - it's pretty central to the paper, but not inspected at all. What is the contents of step 12? How confident are we it's not just a quality control issue?

Cite this project

@misc{gallego2026different,
  title = {{Different Reviewers, Same Blind Spot: Replicating and Specializing Monitors on a Real AI Containment Failure}},
  author = {Sofia Gallego and Muhammad Shoaib Altaf and Devesh Sawant and Xin Li},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/different-reviewers-same-blind-spot-replicating-and-specializing-monitors-on-a-real-ai-containment-failure-rj68}},
  url = {https://apartresearch.com/sprints/projects/different-reviewers-same-blind-spot-replicating-and-specializing-monitors-on-a-real-ai-containment-failure-rj68}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026