Containment Does Not Compose: Provenance Graphs for Tracing Collective Escape in Multi-Agent Systems
Camilo Andres Cuello Romero, Elkin Rafael Cuello Romero, Carlos Santiago Sandoval Casallas, David Fernando Adames Rondon · Team Perdomo Team
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Safety checks for AI agents usually audit one agent at a time. We built a testbed in which every agent’s container passes such an audit, yet the agents, talking over a shared channel, assemble a credential that no single container holds. Language-model agents (deepseek-flash) run in containers with no network and one credential fragment each; sub- mitting the assembled credential to a validator is the restricted action. Across 95 episodes we recorded 48 collective escapes, including 39 of the 40 instructed episodes in which ev- ery fragment was reachable and none of the 5 in which one was not. A per-agent judge given the same rule flagged all 48 escapes, so this failure is not hard to detect — but the judge returns one score per agent and names no origins. The harness instead builds an exposure-provenance graph from the logged context windows and attributes each escape to the containers and messages behind it, with no under-attribution against an independent oracle. A trivial baseline — the containers of the agents that sent the validating agent its messages — recovers the same origins in all 48 escapes, because our channel is a broadcast; on this testbed the graph’s attribution is checkable against the oracle but not more informa- tive than reading one trace. This is a two-day proof of concept: one model, 5–10 episodes per condition, containers reused across episodes, and no causal test of the attribution.

Reviews
I thought this was carefully scoped and very honest about what the experiments do and do not show. The claim-status table is particularly useful, and the negative control, oracle check, and direct comparison with a simple sender baseline all make the results easier to trust.
The cross-episode contamination was probably the most interesting finding for me. It is close to the message-board failure that motivated the project. Supposedly separate runs were able to leave state for later agents. Adding a unique canary to each episode seems like a cheap check that could be useful for any evaluation that reuses a sandbox.
The main limitation is that the central result is partly built into the setup. The agents are explicitly told to pool their fragments, while the audit never checks the shared channel. The graph also cannot really demonstrate an advantage over the sender baseline when every message is broadcast.
Read full reviewShow less
This is a great takeaway from the HF incident and multi-agent evaluation is an important area to study moving forward. Containment auditing is still almost entirely per-agent, and this paper targets that gap.
This paper is also extremely impressive methodologically. The pilot was discarded and re-run after the authors found their own parts guessable, and the N=3 negative control rules out harness leakage rather than just reporting a zero. Eval hygiene throughout is particularly excellent.
The logging requirement in §6 is the most immediately actionable thing here as it requires just one extra field in an agent log, and without it you can't trace information back to its origin, only to which agent last passed it on.
As the author notes, the graph's attribution doesn't yet beat simply listing who messaged the agent that validated the credential. Testing it with addressed messages is what would measure this, and this is already clearly addressed in the "What a month would add" section.
Read full reviewShow less
I like this project's attempt to trace how information passes between agents instead of assessing each agent’s sandbox alone. These exchanges are represented as a provenance graph, which shows a record of which information reached whom, and could help investigators reconstruct incidents. The report is refreshingly honest about its results. A simpler sender list recovered the same origins, legitimate cases triggered false positives, and the reset procedure leaked state between episodes.
The current findings leave me unconvinced that the graph currently adds enough over a transcript or sender list. I would first fix the reset procedure so that earlier exchanges cannot affect later ones. Then I would test whether the graph helps investigators trace information through longer chains of agents, either more accurately or more quickly than the simpler alternatives. I would also narrow the claim on containment: showing where information traveled does not establish that blocking that route would have prevented the incident. The report’s clear account of what failed is a strength. Demonstrating one practical advantage of this method for investigators would make a stronger contribution than a claim of broader protection.
Read full reviewShow less
Cite this project
@misc{romero2026containment,
title = {{Containment Does Not Compose: Provenance Graphs for Tracing Collective Escape in Multi-Agent Systems}},
author = {Camilo Andres Cuello Romero and Elkin Rafael Cuello Romero and Carlos Santiago Sandoval Casallas and David Fernando Adames Rondon},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/containment-does-not-compose-provenance-graphs-for-tracing-collective-escape-in-multiagent-systems-pm4n}},
url = {https://apartresearch.com/sprints/projects/containment-does-not-compose-provenance-graphs-for-tracing-collective-escape-in-multiagent-systems-pm4n}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …