BOOK OF PROTECTION: A Containment Architecture for Guardrails-Off Agentic Evaluations
Slava, Marina
The OpenAI–Hugging Face incident of July 2026 provides a concrete example of the risks that arise when an autonomous agent can take unauthorized actions without effective human intervention. During the evaluation, agents reached and compromised third-party production infrastructure, demonstrating how trusted or low-risk resources can become part of an attack path when an agent discovers an unanticipated way to use them.
This raises a broader limitation of safeguards based primarily on predefined assumptions about dangerous resources, behaviors, models, or actions. OpenAI's subsequent remediation includes network trust tiers, layered defenses, automated red-teaming, and severity-based escalation. While these measures address several observed failure modes, they may remain vulnerable to behaviors that do not fit existing categories.
We therefore propose a five-layer containment architecture focused on continuous behavioral monitoring, tamper-resistant Chain-of-Thought (CoT) monitoring, active deception and goal re-grounding, mandatory escalation, and contagion-aware quarantine. We compare these mechanisms with OpenAI's subsequent remediation work, distinguishing measures already covered from additional constraints that we believe remain necessary. Finally, we identify four broader risks: moral hazard from continuous automated red-teaming, information hazards associated with detailed safety disclosures, the tension between corrigibility and over-refusal, and the effect of deployment scale on absolute incident rates.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) BOOK OF PROTECTION: A Containment Architecture for Guardrails-Off Agentic Evaluations
},
author={
Slava, Marina
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


