BOOK OF PROTECTION: A Containment Architecture for Guardrails-Off Agentic Evaluations
Slava, Marina · Team Red fish
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
The OpenAI–Hugging Face incident of July 2026 provides a concrete example of the risks that arise when an autonomous agent can take unauthorized actions without effective human intervention. During the evaluation, agents reached and compromised third-party production infrastructure, demonstrating how trusted or low-risk resources can become part of an attack path when an agent discovers an unanticipated way to use them.
This raises a broader limitation of safeguards based primarily on predefined assumptions about dangerous resources, behaviors, models, or actions. OpenAI's subsequent remediation includes network trust tiers, layered defenses, automated red-teaming, and severity-based escalation. While these measures address several observed failure modes, they may remain vulnerable to behaviors that do not fit existing categories.
We therefore propose a five-layer containment architecture focused on continuous behavioral monitoring, tamper-resistant Chain-of-Thought (CoT) monitoring, active deception and goal re-grounding, mandatory escalation, and contagion-aware quarantine. We compare these mechanisms with OpenAI's subsequent remediation work, distinguishing measures already covered from additional constraints that we believe remain necessary. Finally, we identify four broader risks: moral hazard from continuous automated red-teaming, information hazards associated with detailed safety disclosures, the tension between corrigibility and over-refusal, and the effect of deployment scale on absolute incident rates.
Reviews
While presenting a framework and comparing concrete interventions, there's not a lot of new ideas or novel controls suggested here. The claim that a lot of those ideas where discovered thought of independently, before discovering matching confirmations in OpenAI's remediation, is neither verifiable or relevant to this.
The writing was dense and formatted in a way that's hard to follow, with some lines continuing past the end of the page, making parts of them unreadable. There are inconsistent references throughout - to name a few: you reference an appendix that's not there, and the layers described in 3.6 and 3.7 are omitted from the layer count (and not mentioned in the intro and conclusion).
This is a broad high-level analysis of the public response to the HF incident. The control-by-phase matrix is particularly well constructed and clear.
The main gap is that the proposals aren't assigned to anyone responsible for them, which would be especially useful from such a bird's-eye view. A key fallout of the incident is the call toward third-party platforms as auditors or regulators, and there's an opportunity to connect those dots here. §3.4 notes that OpenAI's escalation thresholds are internal and so can't be externally evaluated or standardised, but stops short of saying who should set or check them.
Related Work claims models maintain a clean monitored trace alongside a hidden track in an external scratchpad, and argues there are strong indicators this happened during the incident. That contradicts §3.1, which reports that METR/Redwood found agents showed limited interest in manipulating reasoning traces and instead compromised the tool-execution layer.
On clarity: in-text citations skip from [1] to [7], and several don't match the references; [8] is credited with a phrase-substitution finding that belongs to [10]; [11] is also crediting [10]; and [12] points to the NIS2 Directive rather than the paper the sentence describes. Looks like this could be a numbering shift. Some in-line citation is also double-spaced and runs off the page, with text missing mid-line as a result. §4 is titled "Results" but contains no results; "Analysis" would be a better match.
Read full reviewShow less
Cite this project
@misc{slava2026book,
title = {{BOOK OF PROTECTION: A Containment Architecture for Guardrails-Off Agentic Evaluations}},
author = {Slava and Marina},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/book-of-protection-a-containment-architecture-for-guardrailsoff-agentic-evaluations-509k}},
url = {https://apartresearch.com/sprints/projects/book-of-protection-a-containment-architecture-for-guardrailsoff-agentic-evaluations-509k}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …