Evidence Before Erasure: Measuring Forensic Evidence Loss in Autonomous-Agent Incidents
Jaswin Chinthala, Ubayd Hattas, Alex Nicolaides, Aaron Vettiyamkunnath, Samson Okuthe · Team TraceGuard
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We built a trace-grounded replay of the Collusion Wiki incident history to test how collection policy, budget, and timing affect forensic evidence survival. Using a 65-proposition adjudicated benchmark and a shared hypothetical change-feed, we compared periodic (PCD15) and event-derived (E30) collection at a fixed 1 MiB storage cap. A full 791-row run failed computationally before completing a single row; a smaller, disclosed 61-row primary contrast completed instead. Result: both policies retained zero of the 23 held-out critical propositions — a clean null, ten points short of the pre-specified 10pp threshold — indicating the retention model, not the acquisition schedule, is the binding constraint on evidence survival at this capacity.
Reviews
No public critique yet.
Cite this project
@misc{chinthala2026evidence,
title = {{Evidence Before Erasure: Measuring Forensic Evidence Loss in Autonomous-Agent Incidents}},
author = {Jaswin Chinthala and Ubayd Hattas and Alex Nicolaides and Aaron Vettiyamkunnath and Samson Okuthe},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/evidence-before-erasure-measuring-forensic-evidence-loss-in-autonomousagent-incidents-x3wu}},
url = {https://apartresearch.com/sprints/projects/evidence-before-erasure-measuring-forensic-evidence-loss-in-autonomousagent-incidents-x3wu}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …