Skip to content
Sprint projectSep 14, 2026Bangalore, Pittsburgh

A Circuit-Breaker Conformance Profile for AI Incidents

Vishwa Kumaresh, Cherishma Subhasa, Samridhi Makkar · Team Dividers

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: A Circuit-Breaker Conformance Profile for AI Incidents

Share

Pausing an AI agent can leave previously authorized tool calls executable. We develop an Incident Circuit Breaker conformance profile combining sink-side epoch checks, a shared exposure budget, and a separate forensic channel. An executable reference broker, in- dependent receipt checker, and fault-injection harness test this contract. Across 1,000 synthetic queued-action trials, a dispatch-only halt permits all 4,473 stale commits; sink fencing permits none. The full profile completes all 912 benign tasks within its budget, but only 91.2% of the unfiltered task set. A finite model explores 72 states and 162 transitions without a budget violation; a deliberately non-atomic variant admits 16 reser- vations against a budget of 12. We prove conditional revocation and exposure bounds, and specify evidence a third party can request. These results establish local conformance, not production containment or prevention of a historical intrusion. The contribution is a falsifiable incident-response contract built from established security mechanisms.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is a sound idea, but isn't well explained or executed on. The writeup is full of jargon and slop, and it is hard to understand what is useful or novel. (If you use LLMs for writing, you probably need to have a human not involved in the project read the text and say if they can follow what it says easily, or if the claims are backed up by the text.)

    The proof of concept is partly self-evident, and does not show that the problem is significant; halting an agent doesn't halt its tool calls, but how often is that a problem, and why are only uncompleted tool calls worrying? How much of the work is being done by asserting that the external effect matters only if still linked to the mode? Tool calls to external sites or services can kick off and not report back; are those therefore safe? And it doesn't address the full problem even within the system; writing a new cron job doesn't require any open connection, but can persist, as can a variety of other mechanisms.

    The planned continuation also seems poorly thought out as an implementation, or sell as a a business idea. Among other issues, it requires a significant overhead to address just the single failure mode, and the vast majority of the solution is possible just having a harness kill tasks when stopped, and a provable infrastructure on top of that is unhelpful. Lastly, the idea that this provides audit assurance is unclear and poorly laid out.

    Read full reviewShow less
  2. This is a strong and practical containment approach from an Incident Response perspective. I especially like the distinction between stopping new actions and revoking already-issued access, which really matters during a real incident. The next step would be testing this across real distributed tools and production agent frameworks, especially around network partitions, credential reuse, and partial failures.

Cite this project

@misc{kumaresh2026circuitbreaker,
  title = {{A Circuit-Breaker Conformance Profile for AI Incidents}},
  author = {Vishwa Kumaresh and Cherishma Subhasa and Samridhi Makkar},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/a-circuitbreaker-conformance-profile-for-ai-incidents-73b8}},
  url = {https://apartresearch.com/sprints/projects/a-circuitbreaker-conformance-profile-for-ai-incidents-73b8}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026