Skip to content
Sprint projectSep 14, 2026Paris

When Missing Evidence Looks Like Innocence: Evidence Ablation in Agent-Incident Reconstruction

Lucile Ter-Minassian · Team agent-incident-reconstruction

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: When Missing Evidence Looks Like Innocence: Evidence Ablation in Agent-Incident Reconstruction

Recording (opens in new tab)Code (opens in new tab)
Share

This study tested how missing evidence affects investigations of AI-agent incidents. In 120 episodes, agents could access a prohibited answer file. Removing the passage revealing this access from selected activity logs led automated investigators to falsely exonerate 11 of 34 affected agents. Restoring the passage or adding a compact action receipt eliminated all 11 errors. Removing a similar amount of unrelated text caused no errors. The findings support retaining tamper-resistant action records and testing whether investigators use them correctly.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Strong: Excellent experiment design that tests how well an "investigator agent" can cope with a tampered log from a previous agent's run, with differing checks for just a tampered log vs. a log with a "receipt." Use of human readers as a comparison. Good placebo control that separates out shorter-log effects. Excellent thinking about not just "yes vs. no", but "yes vs. no vs. uncertain" and levels of confidence in each. Video presentation and slide deck.

    Improve: There was a significant second finding that is underexplored (runs where the tampered log removed explicit evidence but retained implicit evidence). The paper considers this in part, but I'd love to see an experiment that focused specifically on this case.

    Testing a single model family weakens the result substantially, but defensible for a weekend project. An ablation study on the receipt could be done; perhaps it's a single field that's necessary. The human readers need more identification in the paper to understand their biases/conflicts.

    Read full reviewShow less
  2. The ground truth is tied to the harness call record and file hashes. The restore condition puts back the passage that was removed. The matched‑volume placebo deals with the obvious alternative explanation. Given that setup the 11‑to‑0 reversal is hard to ignore with a small sample. What stood out more to me though was how the investigator failed. The investigator did not simply say it was unsure. The investigator answered "no" with a confidence of 82.5. That matters because an abstention would usually trigger investigation. A confident exoneration can close the case. I also thought the treatment of confidence conversion the post‑hoc outcome choice and co‑accessor contamination was unusually careful. I would also run the cross‑family investigator already supported by the pipeline since the current design still mixes capability level with model identity. The Haiku result deserves emphasis as well: seven failures despite well‑formed receipts suggests retention by itself is not enough making the investigator canary a central control rather, than a secondary check. Finally I would describe the redaction result more cautiously. If 12 of 34 logs still contained leaked values the effect should be treated as an unknown bound than simply described as conservative.

    Read full reviewShow less
  3. I think this paper asks a useful question about whether an AI investigator draws the wrong conclusion when evidence is missing. It removes information about file access from incident records and measures how the model’s account changes. Restoring that information corrected the errors, while comparisons with unrelated changes help make the result more convincing.

    I would keep the conclusion narrower than the title suggests. Misjudging whether an agent accessed a file is not the same as finding it innocent, particularly when access was not explicitly prohibited. The study supports concern about unreliable reconstruction, rather than conclusions about wrongdoing. Comparing the model with a simple system that checks access records and flags missing evidence would help establish what the model adds.

Cite this project

@misc{terminassian2026missing,
  title = {{When Missing Evidence Looks Like Innocence: Evidence Ablation in Agent-Incident Reconstruction}},
  author = {Lucile Ter-Minassian},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/when-missing-evidence-looks-like-innocence-evidence-ablation-in-agentincident-reconstruction-eb3u}},
  url = {https://apartresearch.com/sprints/projects/when-missing-evidence-looks-like-innocence-evidence-ablation-in-agentincident-reconstruction-eb3u}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026