Containment Without Cooperation: A Verifiable Standard for Third-Party Platforms Hosting AI Evaluation Traffic
Hadiza Yusuf · Team Pearl
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Containment standards that a lab checks against its own telemetry face a structural limit: a signed record only proves a known checker ran, not that the underlying data was not curated. This paper relocates verification to the party with no incentive to look away: the third-party platform whose infrastructure an AI agent touches during an evaluation. We define an AI Evaluation Identity (AEI), a credential a lab declares at provisioning time as belonging to an autonomous agent, and specify four containment invariants checkable entirely from a third- party platform’s own access, authorization, and volume logs. Two invariants are preventive and require a declared AEI; two are behavioral and require no disclosure at all. Applied to the July 2026 OpenAI/Hugging Face incident, only the two disclosure-independent invariants fire, because no AEI was ever declared for the campaign. This asymmetry is our central finding: a containment standard’s usefulness against an uncooperative actor depends entirely on which of its checks require that actor’s cooperation. We provide a reference monitor, an illustrative event trace built from the incident’s public record, and a detection-latency estimate showing a volume-based check would plausibly have flagged the campaign within its first day, versus the roughly 3.5 days it actually took.
Reviews
This paper offers a useful framework for platform-side containment checks, distinguishing those that require declared agent identities from those that can operate without developer cooperation. Its illustrative incident analysis makes that distinction concrete, while the claims about earlier detection and demonstrated rule violations need stronger support.
Strength: The distinction between checks requiring developer cooperation and checks an affected platform can perform independently is useful. It directs attention to defenses that remain available when an agent's operator has not declared its activity.
Recommendation:
- Clarify what the standard establishes. Detecting suspicious behavior and demonstrating that a platform failed its response obligations are different outcomes. For example, a volume spike followed by the required throttling satisfies the response rule, even though suspicious traffic occurred. Organize the standard around the observation, required action, and evidence of compliance so its verdicts have a clear security meaning.
- Evaluate the approach from the platform's actual vantage point. The early-detection claim currently uses campaign totals for a rule applied to individual identities. Test what the platform could observe and act on at the time, including legitimate high-volume activity. This would establish whether the rules support useful intervention without excessive disruption.
Read full reviewShow less
This is a thoughtful, well motivated piece of work on how to keep AI systems contained. Its best insight is that checking whether an AI has been kept in check depends partly on who holds the evidence, not just on how carefully the rules are written. The authors suggest that the outside platform an AI actually interacts with is a better source of proof than the AI lab's own records, and that's a useful alternative to simply trusting the lab.
The idea of splitting the rules into two groups is especially helpful. Some rules only work if the AI openly declares that it's an AI being tested. Others cover behavior that can be spotted whether or not the AI cooperates. Spelling this out makes the assumptions behind the standard clear. In the test, the two rules that don't depend on the AI's cooperation triggered, and the two that do were marked "not applicable." That's an informative outcome, not a flaw.
The researcher also built a working example and used it to reconstruct a real incident, which makes the proposal concrete. They were careful not to oversell the 500 actions per day limit, making clear it's just an illustrative number and not a proven real world setting.
The main weakness is that the framework was only checked against the same incident it was built from. To show it works more broadly, the researcher should test it on other, independent containment incidents, tune the behavioural limits using normal, legitimate platform traffic, and try it in a real deployment where AI declare their identities. Overall, this is a very promising framework with a valuable central idea and a clear path to stronger evidence.
Read full reviewShow less
I think this project addresses an important gap because agents can affect platforms that had no involvement in the evaluation. Those platforms need ways to investigate the activity without depending on frontier lab’s account of what happened. The paper proposes monitoring whether agents exceed their permissions, access another customer’s resources, switch between suspicious accounts, or make too many requests.
I am less convinced that the implementation demonstrates reliable detection as this mainly relied on public incident reports and several parts are told which activity is suspicious instead of identifying it themselves. The framework may be useful, but the example does not yet show whether it can distinguish violations from ordinary use. Testing a specific proposed feature against realistic traffic would provide stronger evidence without requiring full deployment. For example I would want to see whether its request volume alerts help identify agents without also flagging ordinary users making legitimate requests.
Read full reviewShow less
Cite this project
@misc{yusuf2026containment,
title = {{Containment Without Cooperation: A Verifiable Standard for Third-Party Platforms Hosting AI Evaluation Traffic}},
author = {Hadiza Yusuf},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/containment-without-cooperation-a-verifiable-standard-for-thirdparty-platforms-hosting-ai-evaluation-traffic-9sz9}},
url = {https://apartresearch.com/sprints/projects/containment-without-cooperation-a-verifiable-standard-for-thirdparty-platforms-hosting-ai-evaluation-traffic-9sz9}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …