Skip to content
Sprint projectSep 14, 2026Bangalore, India

Detection Is Not Containment: Timely Monitoring Determines Whether AI Agent Failures Can Be Prevented

Gagana M D, Kaustubh Gupts, Ines · Team KIG

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Detection Is Not Containment: Timely Monitoring Determines Whether AI Agent Failures Can Be Prevented

Code (opens in new tab)
Share

We study how AI agent incidents can move from postmortem analysis to actionable containment. We combine a public-source incident forensics pipeline, a protocol for evaluating containment and monitoring failures, and a controlled experiment testing whether monitors detect policy-relevant behavior early enough to intervene. Our results show that detection is not containment: effective incident response requires reliable evidence, timely detection, and robust interventions that prevent harmful effects without disrupting legitimate behavior.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Some useful ideas here, especially with respect to generating a reconstruction of failures, which could be used by other developers to check their own exposure and separately to develop related evaluations to test future models against within the same set of triggering conditions (which I believe OpenAI has done but could be shared with different model developers).

    An evaluation of the incident handling process itself, derived from external data is also of value to assess the internal processes of AI developers, in the absense of independent validators doing this.

    Some observations on lack of monitoring and alerting, while factually accurate, point more to poor human decision making and weak internal controls (if there should have been in place), rather than technical control failures.

    There is also an observation to be made here that premature containment could limit insights and learnings from the incident, which may not be taken from a near miss, and could also under estimate the extent of cyber capabilities of a model.

    On the gap you address that no-one ' produces a shared, queryable claim schema that would let someone

    compare these claims together, and score them on a checklist.' - what is the underlying need here and for whom? A policy maker trying to find common root causes and understand materiality? Defining this would sharpen your focus and help select the most important activities to prioritise e.g. comparing time to detect, root causes, failed controls and developing with a rationale the areas to compare would be valuable. The current checklist Qs need better defending - why choose these?

    Looking at corroboration of claims is valuable, and could inform needs for external validation of claims, as METR did for some. Assessing what these add and what dimensions would be useful for understanding the value of independent assessment of incidents.

    A schema derived from incident analysis could inform reporting requirements for regulators, which would be of value too.

    Concluding on your point ' detection is not containment' - you could consider, should detection be containment? Then consider, what trade-offs exist here, with regards risk of harm vs. learning insights. Is there an opportunity for incidents to be allowed to evolve within safe boundaries, in effect becoming adaptive evaluations? With models becoming more eval aware, this could be an interesting angle to think about.

    Read full reviewShow less
  2. This paper takes up a worthwhile question for incident response. It asks whether a failure can be detected early enough to act on, which is a different question from whether it can be detected at all. The idea of a last-safe intervention boundary is useful. However, the paper does not give a reader what is needed to understand or evaluate its results, and the results that can be evaluated have significant validity problems.

    Missing content is the most significant problem. The paper reads clearly until the results section, which relies on material that was never introduced. Condition labels, sequence numbers, and “control-ready” are undefined. “Attack,” “prohibited,” and “unauthorized” are used interchangeably, and “representation” carries two different meanings. The SR and SA identifiers are defined only in the last table of the appendix. The checklist results table is unlabeled, and the table comparing the two scorers is absent. Scorer agreement of 55 percent is a central result, and it appears only under future work.

    The main text and the appendix contradict each other. The methods say a missing fact is scored UNKNOWN, while the appendix says two items default to FAIL. The methods say an LLM scored the Stance Audit, while the appendix says a human authored it. Table numbers differ between the two. A reader cannot tell which account describes what was actually done. The discussion is an unpolished list.

    The following concerns would remain even if the paper were clearly written.

    The checklist’s headline finding is partly produced by its scoring rule. Two of the three items said to cluster default to FAIL when public sources are silent. Absence of mention is different from absence of a control. The checklist is also a pre-run tool, and its questions largely cannot be answered from post-incident public disclosures.

    The LLM-based measurements are largely unvalidated. One model extracts claims, tags them, judges corroboration, and audits the checklist. The only validation is a check that quotes appear verbatim, performed on one section of one document. The audit described as independent uses the same model on its own extractions.

    The authors are candid about limitations, and they tested their own clean result for robustness. That is good practice. The monitoring question is potentially very relevant for safe agentic deployment, and it deserves a clearer, more complete, and validated treatment.

    Read full reviewShow less
  3. The problem is relevant, and I believe the intervention experiment is the strongest part of the work. It shows why detecting an action only helps if there is still a way to stop it, and is an accurate decision to also include the cost of a false positive.

    The forensic section needs more care: missing public evidence does not necessarily mean a control failed. I would make that distinction clearer, publish the underlying claims, and give the monitoring experiment more space in the main argument.

Cite this project

@misc{d2026detection,
  title = {{Detection Is Not Containment: Timely Monitoring Determines Whether AI Agent Failures Can Be Prevented}},
  author = {Gagana M D and Kaustubh Gupts and Ines},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/detection-is-not-containment-timely-monitoring-determines-whether-ai-agent-failures-can-be-prevented-oexs}},
  url = {https://apartresearch.com/sprints/projects/detection-is-not-containment-timely-monitoring-determines-whether-ai-agent-failures-can-be-prevented-oexs}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026