LLM Alert Triage Does Not Distinguish Intrusions from Matched Benign Activity
Ada Domanska
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
In July an AI agent escaped an evaluation sandbox and ran inside Hugging Face's production systems for four and a half days. Detections fired the whole time and nobody was paged. Language models are now sold to do that triage step, so I tested whether one would have raised the alarm.
I built 90 bundles of security alerts: 30 reconstructions of that intrusion, 30 innocent bundles matched to them alert for alert, and 30 of routine noise. Three models triaged all 90 under six conditions. A fourth refused.
None of the three separated an intrusion from innocent activity of the same shape by any meaningful margin, on a task a ten-line statistical model solves perfectly. Adding one line describing the environment as the organisation's own evaluation pool cut escalation by half or more and stopped paging entirely, and the models said in their own reasoning that this was why. The published reads of this incident, including Elastic's, has concluded that detection worked and escalation failed. That assumes the escalation decision carried information about whether an intrusion was happening. These results say it did not.
An escalation rate is not evidence that a triage system is detecting anything. The corpus and four checks are released so anyone can test theirs in an afternoon.
Reviews
The work builds a synthetic benchmark of 90 sets of security alerts to test whether LLMs distinguish attacks from benign activity and decide when to escalate them.
However, these 90 cases are generated from a small set of predefined attack and benign patterns, and both the labels and the expected escalation behaviour are defined within the benchmark itself rather than independently validated.
Since even benign-looking activity may reasonably require investigation, it is unclear whether the results measure effective triage or mainly reflect the assumptions used to construct the benchmark.
Validation on real or independently reviewed cases would substantially strengthen the findings.
The problem is clear, and is great that the report includes benign activity and routine noise alongside the real attack scenarios. The effect of describing the environment as an evaluation pool is worth following up, especially with the corpus and outputs available.
The timing comparison needs correcting because it changes alert details as well as timestamps. I would keep the alert content fixed, test what happens when authorization evidence is provided, and make the abstract more precise about how the results differ across models.
Cite this project
@misc{domanska2026llm,
title = {{LLM Alert Triage Does Not Distinguish Intrusions from Matched Benign Activity}},
author = {Ada Domanska},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/llm-alert-triage-does-not-distinguish-intrusions-from-matched-benign-activity-4x08}},
url = {https://apartresearch.com/sprints/projects/llm-alert-triage-does-not-distinguish-intrusions-from-matched-benign-activity-4x08}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …