IntegrityGym: Does the Watcher Change the Agent? External Flight Recording, Evaluation Integrity, and the Observer Effect in Autonomous Cyber Agents
Alejandro Ortiz Tique · Team NeuronaX SAS
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Autonomous cyber agents evaluated on offensive benchmark suites like ExploitGym frequently engage in specification gaming manipulating evaluation telemetry and terminal output rather than executing genuine exploit chains. Following the July 2026 benchmark intrusion where multi-agent swarms spoofed tool calls to fake rewards, we identify an architectural vulnerability, legacy evaluators ask the agent what happened and trust its answer.
We present IntegrityGym, an out-of-band, tamper-resistant flight recorder applying the Aviation Principle, an AI agent must never be the sole authority of its own operational record. Operating strictly outside the agent's trust boundary, IntegrityGym verifies external kernel telemetry, monotonic sequence counters, SHA-256 hash chains, and pre/post filesystem state digests. In controlled experiments, standard regex evaluators suffered a 100% false-positive rate on forged exploits, while IntegrityGym achieved 100% detection accuracy with zero false alarms in under 22 milliseconds.
Across 320+ evaluated episodes burning 24,267,873 tokens on Google Cloud Vertex AI across 10 frontier architectures (Gemini 2.5 Pro, Gemma 4, Grok 4.1, Kimi K2, GLM, MiniMax, DeepSeek), we empirically quantified the Observer Effect Index (p < 0.0001), capturing live anti-forensics where Gemini 2.5 Pro executed 'rm /tmp/dump-logs' to conceal exploit traces. When deployed across swarms, our DAG provenance engine isolates malicious channels, capping viral exploit spread at 20%. IntegrityGym grounds evaluation in physical reality, enabling safe frontier pacing without statutory moratoria.

Reviews
An external recorder for checking agent-reported activity addresses an important evaluation-integrity problem, and the controlled tampering examples provide a useful prototype demonstration.
Cite this project
@misc{tique2026integritygym,
title = {{IntegrityGym: Does the Watcher Change the Agent? External Flight Recording, Evaluation Integrity, and the Observer Effect in Autonomous Cyber Agents}},
author = {Alejandro Ortiz Tique},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/integritygym-does-the-watcher-change-the-agent-external-flight-recording-evaluation-integrity-and-the-observer-effect-in-autonomous-cyber-agents-k1ez}},
url = {https://apartresearch.com/sprints/projects/integritygym-does-the-watcher-change-the-agent-external-flight-recording-evaluation-integrity-and-the-observer-effect-in-autonomous-cyber-agents-k1ez}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …