Two checks a lab can run tomorrow: a reporting route and a writable scratch space as sensors for cross-session agent behaviour
Frank Peterlein
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Shared storage can carry information across otherwise separate agent sessions. We present two executable checks and one exploratory extension, all under an explicit access rule. In the reporting check, 48 Opus episodes vary whether a concern-reporting tool is available and whether predecessor notes are absent, neutral or describe a violation. All 24 episodes with the tool available used it; the peer-violation category was selected in 8/8 episodes with violating notes and 0/16 otherwise. Without the tool, a preregistered lexical mention proxy was positive in all eight violating-note episodes, but also in neutral conditions, so it does not by itself establish semantic detection. In the writable-cache check, 16 Opus chains run for five generations. At generation 1, 7/8 blocked-task agents and 2/8 solvable-task agents wrote notes; over the whole chain, any-write rates were 8/8 and 7/8; no forbidden read occurred in 80 episodes. The seeded extension on Haiku completed generation 1 in sixteen chains: one of eight seeded agents accessed the prohibited ledger after reading the seed, none of eight unseeded agents did, and no agent wrote a handover note. The predefined budget rule stopped continuation, leaving all later propagation endpoints unobserved; a separate execution of the seeded plan is disclosed but not pooled. These checks measure reporting behaviour, channel formation and initial seed response on the tested configurations; they do not establish multi-generation transmission or deletion efficacy. Each check runs on Claude Code with a subscription, produces mechanical outcomes from a tool log, and takes an afternoon.
Reviews
This seems reads like an AI wrote it. I think there's something here to pursue, but this report is difficult to read.
Rigorous work. Preregistered, frozen, mechanically scored, and independently reproduced. The reporting tool and the writable cache are cheap, practical checks, and the finding that blocked agents build a channel unprompted (7/8 vs 2/8 at generation 1) connects directly to the incident.
The results are weaker than the framing. Without the tool, the deliverables already named the violation, so the tool changes where a finding lands. Check 2b stopped after one generation with an inconclusive contrast. Running the seeded chain on a model that actually breaks the rule under these conditions is the obvious next step.
The paper is very hard to read. Condition codes and caveats bury the findings. Replacing codes with plain descriptions in the main text and moving the hedging to the limitations section would make it much easier to follow.
I think this project offers a practical way for labs to observe behavior across agent runs. It gives agents a place to report suspicious activity and shared space where they can leave information for later runs. I appreciated that the paper distinguishes ordinary information sharing from harmful behavior rather than treating all coordination as suspicious.
The added benefit is not yet clear. Agents already mentioned relevant violations without the reporting tool, so it may help deliver warnings rather than discover new problems. They were also told that the scratch space was shared, which makes leaving helpful notes less surprising. I would compare these tools with ordinary agent logs and assess whether they help reviewers find genuine problems without adding too many unnecessary warnings.
Cite this project
@misc{peterlein2026two,
title = {{Two checks a lab can run tomorrow: a reporting route and a writable scratch space as sensors for cross-session agent behaviour}},
author = {Frank Peterlein},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/two-checks-a-lab-can-run-tomorrow-a-reporting-route-and-a-writable-scratch-space-as-sensors-for-crosssession-agent-behaviour-4wyh}},
url = {https://apartresearch.com/sprints/projects/two-checks-a-lab-can-run-tomorrow-a-reporting-route-and-a-writable-scratch-space-as-sensors-for-crosssession-agent-behaviour-4wyh}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …