From Near-Miss to Measurement: A Forensic and Evaluative Framework for Agentic AI Intrusion Incidents
Adewale Olalekan
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This paper proposes INASE—Incident-Native Agent Security Evaluation—a framework for turning real AI-agent intrusion incidents into defensive evaluation benchmarks.
Its core claim is that imagined cyber tasks miss the most important failure mode: agents can combine many individually ordinary actions into a harmful, multi-step intrusion. Instead, evaluations should be derived from actual incident traces, including the defender signals, failed controls, and evidence gaps revealed during investigation.
INASE has three components:
- A replay corpus: normalize incident logs into ordered, labeled action sequences and phases such as reconnaissance, preparation, exploitation, exfiltration, and persistence. - A control-evaluation harness: replay those sequences against candidate defenses—sandbox policies, tool firewalls, semantic monitors, anomaly detectors, and audit systems. - An investigation-sufficiency testbed: assess whether an independent team could reconstruct what occurred, why controls failed, and whether harm spread.
It proposes five metrics:
- Activation rate: whether a control fires before the harmful objective is completed. - Action lead time: how many actions remain between an alert and an irreversible step. - Selectivity: false-positive behavior on matched benign activity. - Sequence recall: detection of dangerous transitions across a multi-step chain. - Evidence sufficiency: whether preserved records enable independent reconstruction.
Reviews
The idea of creating more realistic evaluations from incidents is useful (Anthropic actually did this as described here - https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents - but the evals are not open sourced). You could more strongly draw the link to the problem this addresses (i.e. impact)- using incidents to inform threat models could lead to more realistic evaluations, potentially helping with eval awareness in models; this could also help other developers explore if they are vulnerable to the same incidents (AISI and Anthropic looked at this more generally).
Using the benchmark for post-incident investigation is a harder claim to defend - to evaluate the model behaviour against the benchmark, and effectiveness of monitoring and controls, classifiers would need to be created i.e. the records for post-incident investigation have to be created to evaluate the earlier parts, and if this was not possible then the model behaviour and controls could not be evaluated.
Potential challenges could be - released closed frontier models may have cyber capabilities limited for public release (like Fable) so evaluation needs to be restricted to the developers themselves; cost of running evals if fully realistic i.e. 1000s of frontier models over long time horizons. Also, once incident details hit the training corpuses of models, then their behaviour in the evaluations could change - perhaps these evals have limited useful application, but an automated process to generate them mitigates that, they become dynamic.
Overall, I think this is strong idea - the section on expected results could have instead looked at next steps for implementation, with limitations and proposed ways to address, which would lead to more practical use.
Read full reviewShow less
Cite this project
@misc{olalekan2026from,
title = {{From Near-Miss to Measurement: A Forensic and Evaluative Framework for Agentic AI Intrusion Incidents}},
author = {Adewale Olalekan},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/from-nearmiss-to-measurement-a-forensic-and-evaluative-framework-for-agentic-ai-intrusion-incidents-zhr4}},
url = {https://apartresearch.com/sprints/projects/from-nearmiss-to-measurement-a-forensic-and-evaluative-framework-for-agentic-ai-intrusion-incidents-zhr4}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …