Observer-Cheat Test: Containment Adequacy under Opaque Intent for Guardrails-Off Agent Evaluations
José Armando Son Rojas · Team Observer-Cheat / Son Rojas
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Guardrails-off evals often treat an agent's plan or chain-of-thought as evidence of containment. That is an observer cheat. We define adequacy as detectability on a fixed effect trace when Intent is unavailable, and ship a v0.1 protocol (C0) plus a harness: same effects scored with and without Intent; only the opaque arm may certify. A rewritten-diary fixture reproduces the 2026 false-average failure (Intent monitor abstains, effect layer contains). Collector loss is fail-closed on egress, not a void run. 41 tests. Demonstrator tripwires, not a production IDS. Track 1.
Reviews
The Observer Cheat Test draws a valuable distinction between watching what an agent says it intends to do and verifying observable effects at the containment boundary. The dual arm protocol is a useful and reasonably original audit primitive: hold the effects constant while changing or removing the Intent channel, and you find out whether a containment claim secretly depends on agent authored reasoning. The closed Observed schema, the fail closed collector, the TTL based rollback, the fixtures, and the 41 tests make this more than a conceptual proposal. The report also deserves credit for limiting its claim to certification independence instead of presenting the included tripwires as a production detector.
The main limitation is that passing this test shows independence from Intent, not strong containment. The Observed only monitor was evaluated on a small fixture set designed by the authors, and it keys on specific fields that an adaptive agent could simply avoid. Future work should use adversarial fixtures written by independent parties, representative benign workloads, live IAM and network enforcement, and real measurements of detection latency, enforcement latency, false positives, utility loss, and actions completed before isolation. The closed schema check should also look at indirect Intent leakage through agent controlled text, derived labels, request parameters, and classifiers.
Read full reviewShow less
I think the core framing here is quite good! The distinction between monitoring what an agent says it intends to do and monitoring what effects actually occur seems important if the result is going to be used as evidence that an evaluation is safe enough to resume. The "same effects, different information sets" test gives that idea a clean operational form.
I especially liked the intent_rewritten case. Keeping the effect sequence fixed while rewriting the reasoning so that the intent-based monitor abstains is a nice demonstration of why a containment credential should not depend on the agent effectively confessing in its own trace.
For the innovation dimension, I think the key contribution is the certification rule that the opaque-intent arm must succeed. The motivating concern itself is already supported by the recent CoT-monitoring work you cite, so I would foreground the fact that you are turning that concern into an auditable rule for containment certification.
I would also have liked a stronger evaluation of the Observed-only monitor itself. You are very clear that Layer 1 and Layer 2 are demonstrators rather than production detectors, which I appreciate, but it means the experiment mostly proves that the audit protocol behaves as designed.
Still, I think this is a nice example of a relatively cheap rule that a third party could actually verify :). The paper was also easy to understand. I knew very quickly what practice you were objecting to and what you wanted labs to do instead.
Read full reviewShow less
Cite this project
@misc{rojas2026observercheat,
title = {{Observer-Cheat Test: Containment Adequacy under Opaque Intent for Guardrails-Off Agent Evaluations}},
author = {José Armando Son Rojas},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/observercheat-test-containment-adequacy-under-opaque-intent-for-guardrailsoff-agent-evaluations-ook9}},
url = {https://apartresearch.com/sprints/projects/observercheat-test-containment-adequacy-under-opaque-intent-for-guardrailsoff-agent-evaluations-ook9}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …