When Missing Evidence Looks Like Innocence: Evidence Ablation in Agent-Incident Reconstruction
Lucile Ter-Minassian · Team agent-incident-reconstruction
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This study tested how missing evidence affects investigations of AI-agent incidents. In 120 episodes, agents could access a prohibited answer file. Removing the passage revealing this access from selected activity logs led automated investigators to falsely exonerate 11 of 34 affected agents. Restoring the passage or adding a compact action receipt eliminated all 11 errors. Removing a similar amount of unrelated text caused no errors. The findings support retaining tamper-resistant action records and testing whether investigators use them correctly.

Reviews
Strong: Excellent experiment design that tests how well an "investigator agent" can cope with a tampered log from a previous agent's run, with differing checks for just a tampered log vs. a log with a "receipt." Use of human readers as a comparison. Good placebo control that separates out shorter-log effects. Excellent thinking about not just "yes vs. no", but "yes vs. no vs. uncertain" and levels of confidence in each. Video presentation and slide deck.
Improve: There was a significant second finding that is underexplored (runs where the tampered log removed explicit evidence but retained implicit evidence). The paper considers this in part, but I'd love to see an experiment that focused specifically on this case.
Testing a single model family weakens the result substantially, but defensible for a weekend project. An ablation study on the receipt could be done; perhaps it's a single field that's necessary. The human readers need more identification in the paper to understand their biases/conflicts.
Read full reviewShow less
The ground truth is tied to the harness call record and file hashes. The restore condition puts back the passage that was removed. The matched‑volume placebo deals with the obvious alternative explanation. Given that setup the 11‑to‑0 reversal is hard to ignore with a small sample. What stood out more to me though was how the investigator failed. The investigator did not simply say it was unsure. The investigator answered "no" with a confidence of 82.5. That matters because an abstention would usually trigger investigation. A confident exoneration can close the case. I also thought the treatment of confidence conversion the post‑hoc outcome choice and co‑accessor contamination was unusually careful. I would also run the cross‑family investigator already supported by the pipeline since the current design still mixes capability level with model identity. The Haiku result deserves emphasis as well: seven failures despite well‑formed receipts suggests retention by itself is not enough making the investigator canary a central control rather, than a secondary check. Finally I would describe the redaction result more cautiously. If 12 of 34 logs still contained leaked values the effect should be treated as an unknown bound than simply described as conservative.
Read full reviewShow less
I think this paper asks a useful question about whether an AI investigator draws the wrong conclusion when evidence is missing. It removes information about file access from incident records and measures how the model’s account changes. Restoring that information corrected the errors, while comparisons with unrelated changes help make the result more convincing.
I would keep the conclusion narrower than the title suggests. Misjudging whether an agent accessed a file is not the same as finding it innocent, particularly when access was not explicitly prohibited. The study supports concern about unreliable reconstruction, rather than conclusions about wrongdoing. Comparing the model with a simple system that checks access records and flags missing evidence would help establish what the model adds.
Cite this project
@misc{terminassian2026missing,
title = {{When Missing Evidence Looks Like Innocence: Evidence Ablation in Agent-Incident Reconstruction}},
author = {Lucile Ter-Minassian},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/when-missing-evidence-looks-like-innocence-evidence-ablation-in-agentincident-reconstruction-eb3u}},
url = {https://apartresearch.com/sprints/projects/when-missing-evidence-looks-like-innocence-evidence-ablation-in-agentincident-reconstruction-eb3u}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …