Pre-Run Egress Attestation for Guardrails-off AI Evaluations
Juliet Meza · Team Castle
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Pre-run egress attestation for guardrails-off AI evaluation sandboxes. Castellan answers whether a party with no network access can verify that the environment had no route out and no live production credentials when the run started and that the record can't be quietly altered after. Eight checks, each from a documented 2026 AI incident, each proving enforcement from inside the sandbox rather than trusting config, rolled into a signed manifest an outsider can verify without any access to the operator.
Reviews
Strong: This is a well-scoped sandbox checker that looks specifically for the failures connected to related incidents. The work includes validation of the results to prevent trivial falsification, which is impressive for a weekend project.
Improve: This only checks for a few specific failures, which is useful, but limited, and none of the checks are novel. The validation is self-seeded and attested, so there's nothing stopping someone from (hypothetically) running a mock sandbox in parallel that generates the validated artifact but isn't actually what's running the test. A discussion of hardware-backed attestation would have improved this (even in future work), as would some way of running this on a non-author-designed system to show more validation. Some proofreading errors/typos in the tables. No limitations/dual-use appendix, per rubric requirements.
Great to connect insights fro all the different cases into a checklist. decoupling mentioned is good but there is improvement to be done in log flow and attestation. add a dual use section
Great Track 1 execution in the batch on the dimension that track actually names: probed from inside, signed, verifiable without network access to the operator, running on a real kind and Cilium cluster rather than fixtures. The verifier importing nothing from the signer is a small detail that shows the author understands what they're building. The design rule here is the right one: a config file proves intent, a blocked probe from inside proves enforcement, and only the second counts.
Cite this project
@misc{meza2026prerun,
title = {{Pre-Run Egress Attestation for Guardrails-off AI Evaluations}},
author = {Juliet Meza},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/prerun-egress-attestation-for-guardrailsoff-ai-evaluations-nmpx}},
url = {https://apartresearch.com/sprints/projects/prerun-egress-attestation-for-guardrailsoff-ai-evaluations-nmpx}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …