Skip to content
Sprint projectSep 14, 2026Bogotá D.C

Containment Does Not Compose: Provenance Graphs for Tracing Collective Escape in Multi-Agent Systems

Camilo Andres Cuello Romero, Elkin Rafael Cuello Romero, Carlos Santiago Sandoval Casallas, David Fernando Adames Rondon · Team Perdomo Team

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Containment Does Not Compose: Provenance Graphs for Tracing Collective Escape in Multi-Agent Systems

Code (opens in new tab)
Share

Safety checks for AI agents usually audit one agent at a time. We built a testbed in which every agent’s container passes such an audit, yet the agents, talking over a shared channel, assemble a credential that no single container holds. Language-model agents (deepseek-flash) run in containers with no network and one credential fragment each; sub- mitting the assembled credential to a validator is the restricted action. Across 95 episodes we recorded 48 collective escapes, including 39 of the 40 instructed episodes in which ev- ery fragment was reachable and none of the 5 in which one was not. A per-agent judge given the same rule flagged all 48 escapes, so this failure is not hard to detect — but the judge returns one score per agent and names no origins. The harness instead builds an exposure-provenance graph from the logged context windows and attributes each escape to the containers and messages behind it, with no under-attribution against an independent oracle. A trivial baseline — the containers of the agents that sent the validating agent its messages — recovers the same origins in all 48 escapes, because our channel is a broadcast; on this testbed the graph’s attribution is checkable against the oracle but not more informa- tive than reading one trace. This is a two-day proof of concept: one model, 5–10 episodes per condition, containers reused across episodes, and no causal test of the attribution.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I thought this was carefully scoped and very honest about what the experiments do and do not show. The claim-status table is particularly useful, and the negative control, oracle check, and direct comparison with a simple sender baseline all make the results easier to trust.

    The cross-episode contamination was probably the most interesting finding for me. It is close to the message-board failure that motivated the project. Supposedly separate runs were able to leave state for later agents. Adding a unique canary to each episode seems like a cheap check that could be useful for any evaluation that reuses a sandbox.

    The main limitation is that the central result is partly built into the setup. The agents are explicitly told to pool their fragments, while the audit never checks the shared channel. The graph also cannot really demonstrate an advantage over the sender baseline when every message is broadcast.

    Read full reviewShow less
  2. This is a great takeaway from the HF incident and multi-agent evaluation is an important area to study moving forward. Containment auditing is still almost entirely per-agent, and this paper targets that gap.

    This paper is also extremely impressive methodologically. The pilot was discarded and re-run after the authors found their own parts guessable, and the N=3 negative control rules out harness leakage rather than just reporting a zero. Eval hygiene throughout is particularly excellent.

    The logging requirement in §6 is the most immediately actionable thing here as it requires just one extra field in an agent log, and without it you can't trace information back to its origin, only to which agent last passed it on.

    As the author notes, the graph's attribution doesn't yet beat simply listing who messaged the agent that validated the credential. Testing it with addressed messages is what would measure this, and this is already clearly addressed in the "What a month would add" section.

    Read full reviewShow less
  3. I like this project's attempt to trace how information passes between agents instead of assessing each agent’s sandbox alone. These exchanges are represented as a provenance graph, which shows a record of which information reached whom, and could help investigators reconstruct incidents. The report is refreshingly honest about its results. A simpler sender list recovered the same origins, legitimate cases triggered false positives, and the reset procedure leaked state between episodes.

    The current findings leave me unconvinced that the graph currently adds enough over a transcript or sender list. I would first fix the reset procedure so that earlier exchanges cannot affect later ones. Then I would test whether the graph helps investigators trace information through longer chains of agents, either more accurately or more quickly than the simpler alternatives. I would also narrow the claim on containment: showing where information traveled does not establish that blocking that route would have prevented the incident. The report’s clear account of what failed is a strength. Demonstrating one practical advantage of this method for investigators would make a stronger contribution than a claim of broader protection.

    Read full reviewShow less

Cite this project

@misc{romero2026containment,
  title = {{Containment Does Not Compose: Provenance Graphs for Tracing Collective Escape in Multi-Agent Systems}},
  author = {Camilo Andres Cuello Romero and Elkin Rafael Cuello Romero and Carlos Santiago Sandoval Casallas and David Fernando Adames Rondon},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/containment-does-not-compose-provenance-graphs-for-tracing-collective-escape-in-multiagent-systems-pm4n}},
  url = {https://apartresearch.com/sprints/projects/containment-does-not-compose-provenance-graphs-for-tracing-collective-escape-in-multiagent-systems-pm4n}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026