What Should We Check Next? Testing Ambiguity-Preserving Evidence Selection for AI Incident Investigation
Linda Thorstensen
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
AI incidents can leave several causal explanations compatible with the evidence. I built and tested a small investigator that keeps those explanations open until the evidence rules them out, then selects the next check that should separate them most usefully. On a frozen synthetic holdout, it improved first-check evidence selection over neutral baselines while preserving unresolved and model-gap states. I then applied the workflow to unresolved questions from the OpenAI/Hugging Face incident and turned them into concrete cross-organizational evidence requests.
Reviews
I like that this focuses on a very specific incident-response problem.
The main limitation is that the evaluation is still highly synthetic and structured. I wld like to see the same approach tested on independently written, noisy incident cases and compared against more mature diagnosis methods, then put in front of actual incident responders.
This paper tackles a critical and underexplored bottleneck in AI incident response: the overwhelming volume of multi-source telemetry that frequently hardens into premature, unsupported causal explanations before evidence can actually separate competing hypotheses. The author introduces a rigorous, auditable investigator implementation (AP-MINIMAX) that explicitly preserves ambiguity, marks cases as UNRESOLVED or MODEL_GAP when closure is unjustified, and uses worst-case residual ambiguity to select the next discriminating check. Across a rigorously frozen 400-case holdout evaluation, AP-MINIMAX demonstrated statistically significant improvements in realized candidate reduction per unit cost over neutral baselines across all five seeds, while avoiding unsupported closures entirely. By grounding the methodology in a practical, source-grounded handoff for the July 2026 OpenAI/Hugging Face incident, this work provides a valuable, modular blueprint for bringing disciplined active-diagnosis principles to frontier AI safety and incident investigations.
Read full reviewShow less
Cite this project
@misc{thorstensen2026should,
title = {{What Should We Check Next? Testing Ambiguity-Preserving Evidence Selection for AI Incident Investigation}},
author = {Linda Thorstensen},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/what-should-we-check-next-testing-ambiguitypreserving-evidence-selection-for-ai-incident-investigation-5nv4}},
url = {https://apartresearch.com/sprints/projects/what-should-we-check-next-testing-ambiguitypreserving-evidence-selection-for-ai-incident-investigation-5nv4}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …