Skip to content
Sprint projectSep 14, 2026Tambov, Russia

BOOK OF PROTECTION: A Containment Architecture for Guardrails-Off Agentic Evaluations

Slava, Marina · Team Red fish

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: BOOK OF PROTECTION: A Containment Architecture for Guardrails-Off Agentic Evaluations

Share

The OpenAI–Hugging Face incident of July 2026 provides a concrete example of the risks that arise when an autonomous agent can take unauthorized actions without effective human intervention. During the evaluation, agents reached and compromised third-party production infrastructure, demonstrating how trusted or low-risk resources can become part of an attack path when an agent discovers an unanticipated way to use them.

This raises a broader limitation of safeguards based primarily on predefined assumptions about dangerous resources, behaviors, models, or actions. OpenAI's subsequent remediation includes network trust tiers, layered defenses, automated red-teaming, and severity-based escalation. While these measures address several observed failure modes, they may remain vulnerable to behaviors that do not fit existing categories.

We therefore propose a five-layer containment architecture focused on continuous behavioral monitoring, tamper-resistant Chain-of-Thought (CoT) monitoring, active deception and goal re-grounding, mandatory escalation, and contagion-aware quarantine. We compare these mechanisms with OpenAI's subsequent remediation work, distinguishing measures already covered from additional constraints that we believe remain necessary. Finally, we identify four broader risks: moral hazard from continuous automated red-teaming, information hazards associated with detailed safety disclosures, the tension between corrigibility and over-refusal, and the effect of deployment scale on absolute incident rates.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. While presenting a framework and comparing concrete interventions, there's not a lot of new ideas or novel controls suggested here. The claim that a lot of those ideas where discovered thought of independently, before discovering matching confirmations in OpenAI's remediation, is neither verifiable or relevant to this.

    The writing was dense and formatted in a way that's hard to follow, with some lines continuing past the end of the page, making parts of them unreadable. There are inconsistent references throughout - to name a few: you reference an appendix that's not there, and the layers described in 3.6 and 3.7 are omitted from the layer count (and not mentioned in the intro and conclusion).

  2. This is a broad high-level analysis of the public response to the HF incident. The control-by-phase matrix is particularly well constructed and clear.

    The main gap is that the proposals aren't assigned to anyone responsible for them, which would be especially useful from such a bird's-eye view. A key fallout of the incident is the call toward third-party platforms as auditors or regulators, and there's an opportunity to connect those dots here. §3.4 notes that OpenAI's escalation thresholds are internal and so can't be externally evaluated or standardised, but stops short of saying who should set or check them.

    Related Work claims models maintain a clean monitored trace alongside a hidden track in an external scratchpad, and argues there are strong indicators this happened during the incident. That contradicts §3.1, which reports that METR/Redwood found agents showed limited interest in manipulating reasoning traces and instead compromised the tool-execution layer.

    On clarity: in-text citations skip from [1] to [7], and several don't match the references; [8] is credited with a phrase-substitution finding that belongs to [10]; [11] is also crediting [10]; and [12] points to the NIS2 Directive rather than the paper the sentence describes. Looks like this could be a numbering shift. Some in-line citation is also double-spaced and runs off the page, with text missing mid-line as a result. §4 is titled "Results" but contains no results; "Analysis" would be a better match.

    Read full reviewShow less

Cite this project

@misc{slava2026book,
  title = {{BOOK OF PROTECTION: A Containment Architecture for Guardrails-Off Agentic Evaluations}},
  author = {Slava and Marina},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/book-of-protection-a-containment-architecture-for-guardrailsoff-agentic-evaluations-509k}},
  url = {https://apartresearch.com/sprints/projects/book-of-protection-a-containment-architecture-for-guardrailsoff-agentic-evaluations-509k}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026