Skip to content
Sprint projectSep 14, 2026Taylor

Containment Without Cooperation: A Verifiable Standard for Third-Party Platforms Hosting AI Evaluation Traffic

Hadiza Yusuf · Team Pearl

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Containment Without Cooperation: A Verifiable Standard for Third-Party Platforms Hosting AI Evaluation Traffic

Presentation

Presentation: Containment Without Cooperation: A Verifiable Standard for Third-Party Platforms Hosting AI Evaluation Traffic

Code (opens in new tab)
Share

Containment standards that a lab checks against its own telemetry face a structural limit: a signed record only proves a known checker ran, not that the underlying data was not curated. This paper relocates verification to the party with no incentive to look away: the third-party platform whose infrastructure an AI agent touches during an evaluation. We define an AI Evaluation Identity (AEI), a credential a lab declares at provisioning time as belonging to an autonomous agent, and specify four containment invariants checkable entirely from a third- party platform’s own access, authorization, and volume logs. Two invariants are preventive and require a declared AEI; two are behavioral and require no disclosure at all. Applied to the July 2026 OpenAI/Hugging Face incident, only the two disclosure-independent invariants fire, because no AEI was ever declared for the campaign. This asymmetry is our central finding: a containment standard’s usefulness against an uncooperative actor depends entirely on which of its checks require that actor’s cooperation. We provide a reference monitor, an illustrative event trace built from the incident’s public record, and a detection-latency estimate showing a volume-based check would plausibly have flagged the campaign within its first day, versus the roughly 3.5 days it actually took.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper offers a useful framework for platform-side containment checks, distinguishing those that require declared agent identities from those that can operate without developer cooperation. Its illustrative incident analysis makes that distinction concrete, while the claims about earlier detection and demonstrated rule violations need stronger support.

    Strength: The distinction between checks requiring developer cooperation and checks an affected platform can perform independently is useful. It directs attention to defenses that remain available when an agent's operator has not declared its activity.

    Recommendation:

    - Clarify what the standard establishes. Detecting suspicious behavior and demonstrating that a platform failed its response obligations are different outcomes. For example, a volume spike followed by the required throttling satisfies the response rule, even though suspicious traffic occurred. Organize the standard around the observation, required action, and evidence of compliance so its verdicts have a clear security meaning.

    - Evaluate the approach from the platform's actual vantage point. The early-detection claim currently uses campaign totals for a rule applied to individual identities. Test what the platform could observe and act on at the time, including legitimate high-volume activity. This would establish whether the rules support useful intervention without excessive disruption.

    Read full reviewShow less
  2. This is a thoughtful, well motivated piece of work on how to keep AI systems contained. Its best insight is that checking whether an AI has been kept in check depends partly on who holds the evidence, not just on how carefully the rules are written. The authors suggest that the outside platform an AI actually interacts with is a better source of proof than the AI lab's own records, and that's a useful alternative to simply trusting the lab.

    The idea of splitting the rules into two groups is especially helpful. Some rules only work if the AI openly declares that it's an AI being tested. Others cover behavior that can be spotted whether or not the AI cooperates. Spelling this out makes the assumptions behind the standard clear. In the test, the two rules that don't depend on the AI's cooperation triggered, and the two that do were marked "not applicable." That's an informative outcome, not a flaw.

    The researcher also built a working example and used it to reconstruct a real incident, which makes the proposal concrete. They were careful not to oversell the 500 actions per day limit, making clear it's just an illustrative number and not a proven real world setting.

    The main weakness is that the framework was only checked against the same incident it was built from. To show it works more broadly, the researcher should test it on other, independent containment incidents, tune the behavioural limits using normal, legitimate platform traffic, and try it in a real deployment where AI declare their identities. Overall, this is a very promising framework with a valuable central idea and a clear path to stronger evidence.

    Read full reviewShow less
  3. I think this project addresses an important gap because agents can affect platforms that had no involvement in the evaluation. Those platforms need ways to investigate the activity without depending on frontier lab’s account of what happened. The paper proposes monitoring whether agents exceed their permissions, access another customer’s resources, switch between suspicious accounts, or make too many requests.

    I am less convinced that the implementation demonstrates reliable detection as this mainly relied on public incident reports and several parts are told which activity is suspicious instead of identifying it themselves. The framework may be useful, but the example does not yet show whether it can distinguish violations from ordinary use. Testing a specific proposed feature against realistic traffic would provide stronger evidence without requiring full deployment. For example I would want to see whether its request volume alerts help identify agents without also flagging ordinary users making legitimate requests.

    Read full reviewShow less

Cite this project

@misc{yusuf2026containment,
  title = {{Containment Without Cooperation: A Verifiable Standard for Third-Party Platforms Hosting AI Evaluation Traffic}},
  author = {Hadiza Yusuf},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/containment-without-cooperation-a-verifiable-standard-for-thirdparty-platforms-hosting-ai-evaluation-traffic-9sz9}},
  url = {https://apartresearch.com/sprints/projects/containment-without-cooperation-a-verifiable-standard-for-thirdparty-platforms-hosting-ai-evaluation-traffic-9sz9}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026