Skip to content
Sprint projectSep 13, 2026Abuja, Nigeria

From Near-Miss to Measurement: A Forensic and Evaluative Framework for Agentic AI Intrusion Incidents

Adewale Olalekan

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: From Near-Miss to Measurement: A Forensic and Evaluative Framework for Agentic AI Intrusion Incidents

Share

This paper proposes INASE—Incident-Native Agent Security Evaluation—a framework for turning real AI-agent intrusion incidents into defensive evaluation benchmarks.

Its core claim is that imagined cyber tasks miss the most important failure mode: agents can combine many individually ordinary actions into a harmful, multi-step intrusion. Instead, evaluations should be derived from actual incident traces, including the defender signals, failed controls, and evidence gaps revealed during investigation.

INASE has three components:

- A replay corpus: normalize incident logs into ordered, labeled action sequences and phases such as reconnaissance, preparation, exploitation, exfiltration, and persistence. - A control-evaluation harness: replay those sequences against candidate defenses—sandbox policies, tool firewalls, semantic monitors, anomaly detectors, and audit systems. - An investigation-sufficiency testbed: assess whether an independent team could reconstruct what occurred, why controls failed, and whether harm spread.

It proposes five metrics:

- Activation rate: whether a control fires before the harmful objective is completed. - Action lead time: how many actions remain between an alert and an irreversible step. - Selectivity: false-positive behavior on matched benign activity. - Sequence recall: detection of dangerous transitions across a multi-step chain. - Evidence sufficiency: whether preserved records enable independent reconstruction.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The idea of creating more realistic evaluations from incidents is useful (Anthropic actually did this as described here - https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents - but the evals are not open sourced). You could more strongly draw the link to the problem this addresses (i.e. impact)- using incidents to inform threat models could lead to more realistic evaluations, potentially helping with eval awareness in models; this could also help other developers explore if they are vulnerable to the same incidents (AISI and Anthropic looked at this more generally).

    Using the benchmark for post-incident investigation is a harder claim to defend - to evaluate the model behaviour against the benchmark, and effectiveness of monitoring and controls, classifiers would need to be created i.e. the records for post-incident investigation have to be created to evaluate the earlier parts, and if this was not possible then the model behaviour and controls could not be evaluated.

    Potential challenges could be - released closed frontier models may have cyber capabilities limited for public release (like Fable) so evaluation needs to be restricted to the developers themselves; cost of running evals if fully realistic i.e. 1000s of frontier models over long time horizons. Also, once incident details hit the training corpuses of models, then their behaviour in the evaluations could change - perhaps these evals have limited useful application, but an automated process to generate them mitigates that, they become dynamic.

    Overall, I think this is strong idea - the section on expected results could have instead looked at next steps for implementation, with limitations and proposed ways to address, which would lead to more practical use.

    Read full reviewShow less

Cite this project

@misc{olalekan2026from,
  title = {{From Near-Miss to Measurement: A Forensic and Evaluative Framework for Agentic AI Intrusion Incidents}},
  author = {Adewale Olalekan},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/from-nearmiss-to-measurement-a-forensic-and-evaluative-framework-for-agentic-ai-intrusion-incidents-zhr4}},
  url = {https://apartresearch.com/sprints/projects/from-nearmiss-to-measurement-a-forensic-and-evaluative-framework-for-agentic-ai-intrusion-incidents-zhr4}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026