The Six Questions: A Composition-Based Declaration for Agent Evaluation Harnesses, Scored Against the 2026 Escapes
Travis Gilly · Team Convergence Working Group
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Every containment control in the 2026 agent escapes was written on a mechanism, no POST, one permitted egress, no inter-agent communication, and each held exactly as written while the denied effect happened anyway, because the granted affordances composed into it. This is a v0.1 standard for evaluation harnesses that run with production safeguards disabled: a six-row declaration in which the operator states which affordances are granted, names the control paired with each, and runs a composition test against every denied effect. Nothing is prohibited, so nothing blocks adoption; the failure condition is an undeclared composition. Every row is answerable from configuration files by a platform engineer, with no model expertise and no access to the lab's network, and each carries a rough cost. Scored against the German wiki swarm and the Hugging Face intrusion, it marks the central denials of both as nominal at configuration time, before either run started.
Reviews
This work highlights an important issue where individual permissions may look safe, but when combined they can create a capability that was supposed to be blocked. I also liked that the proposed declaration is simple and could be practical to use. The main limitation is that most of the testing is based on known incidents. Testing this on new and unseen scenarios, with independent reviewers, would help show how well the approach works in real-world environments.
Calling out that the two failures behind both incidents did not originate from permissions given to the agents is a good observation and takeaway from the incident. The paper is also clear on ownership and responsibility, which makes this a strong response to the incident with a clear path forward.
On the design: the six affordances are derived from the two incidents (the German wiki and OpenAI/Hugging Face) and then scored against them, so the scorings can't fail and the evaluation is circular. Both incidents appear to involve OpenAI agents, so the two scorings show one lab failing twice rather than the rows transferring across deployers
It is important to learn from these incidents and formulate mitigations accordingly, but a held-out test would be stronger: deriving from one incident and scoring the other would have been a better evaluation. Scoring the Anthropic incidents, as Future Work proposes, would also test whether the rows hold beyond one deployer.
Read full reviewShow less
Cite this project
@misc{gilly2026six,
title = {{The Six Questions: A Composition-Based Declaration for Agent Evaluation Harnesses, Scored Against the 2026 Escapes}},
author = {Travis Gilly},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-six-questions-a-compositionbased-declaration-for-agent-evaluation-harnesses-scored-against-the-2026-escapes-3l7l}},
url = {https://apartresearch.com/sprints/projects/the-six-questions-a-compositionbased-declaration-for-agent-evaluation-harnesses-scored-against-the-2026-escapes-3l7l}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …