Skip to content
Sprint projectSep 13, 2026London

LLM Alert Triage Does Not Distinguish Intrusions from Matched Benign Activity

Ada Domanska

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: LLM Alert Triage Does Not Distinguish Intrusions from Matched Benign Activity

Code (opens in new tab)
Share

In July an AI agent escaped an evaluation sandbox and ran inside Hugging Face's production systems for four and a half days. Detections fired the whole time and nobody was paged. Language models are now sold to do that triage step, so I tested whether one would have raised the alarm.

I built 90 bundles of security alerts: 30 reconstructions of that intrusion, 30 innocent bundles matched to them alert for alert, and 30 of routine noise. Three models triaged all 90 under six conditions. A fourth refused.

None of the three separated an intrusion from innocent activity of the same shape by any meaningful margin, on a task a ten-line statistical model solves perfectly. Adding one line describing the environment as the organisation's own evaluation pool cut escalation by half or more and stopped paging entirely, and the models said in their own reasoning that this was why. The published reads of this incident, including Elastic's, has concluded that detection worked and escalation failed. That assumes the escalation decision carried information about whether an intrusion was happening. These results say it did not.

An escalation rate is not evidence that a triage system is detecting anything. The corpus and four checks are released so anyone can test theirs in an afternoon.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The work builds a synthetic benchmark of 90 sets of security alerts to test whether LLMs distinguish attacks from benign activity and decide when to escalate them.

    However, these 90 cases are generated from a small set of predefined attack and benign patterns, and both the labels and the expected escalation behaviour are defined within the benchmark itself rather than independently validated.

    Since even benign-looking activity may reasonably require investigation, it is unclear whether the results measure effective triage or mainly reflect the assumptions used to construct the benchmark.

    Validation on real or independently reviewed cases would substantially strengthen the findings.

  2. The problem is clear, and is great that the report includes benign activity and routine noise alongside the real attack scenarios. The effect of describing the environment as an evaluation pool is worth following up, especially with the corpus and outputs available.

    The timing comparison needs correcting because it changes alert details as well as timestamps. I would keep the alert content fixed, test what happens when authorization evidence is provided, and make the abstract more precise about how the results differ across models.

Cite this project

@misc{domanska2026llm,
  title = {{LLM Alert Triage Does Not Distinguish Intrusions from Matched Benign Activity}},
  author = {Ada Domanska},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/llm-alert-triage-does-not-distinguish-intrusions-from-matched-benign-activity-4x08}},
  url = {https://apartresearch.com/sprints/projects/llm-alert-triage-does-not-distinguish-intrusions-from-matched-benign-activity-4x08}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026