Towards Early Warning for AI Agent Incidents: Monitoring the Dynamics of Risk Trajectories
Yifan Li, John Gia Bao Luc, Rashik Shahjahan, Kai Ting · Team AI response T1
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Recent AI-agent incidents highlight the need for monitoring systems that can detect emerging unsafe behavior during agent execution, rather than only assess failures after they occur. Motivated by this challenge, we study a modular early-warning framework for LLM agents based on the dynamics of their observable trajectories. The framework first maps a trajectory prefix \(H_t\), containing messages, tool calls, observations, and actions, to a numerical risk-relevant score \(S_t\), and then monitors the evolution of the resulting score process over time. We demonstrate the scoring layer on partial real-agent logs using StepGuard to construct action-unsafe proxy score trajectories, and study the temporal layer on seeded controlled trajectories using pointwise level, recent slope (``momentum''), EWMA, and CUSUM statistics. In the constructed unsafe examples, momentum and CUSUM cross their selected thresholds before the designated unsafe action, whereas level and EWMA cross at that action; neither constructed safe trajectory alarms under the selected cutoffs. These results are illustrative rather than calibrated detector comparisons. Overall, this sprint provides a proof of concept for a score-then-monitor approach to agent safety and motivates systematic evaluation of temporal risk dynamics as a basis for earlier warning and human intervention in deployed agent systems.
Reviews
For a weekend sprint, this study is a clearly scoped proof of concept. The core framing of the study: separating the scorer (by turning a prefix trajectory into a risk score) from the temporal monitor (to watch how that score moves), and treating this score as a time series you can run CUSUM/EWMA/momentum on - is the part I liked the most. The authors are also honest about the study being illustrative, not a calibrated early-warning system. The temporal monitoring code is in the repo and reproduces the table; the StepGuard real-log scoring pipeline isn't committed, so that half is not independently reproducible yet, which is understandable for a weekend and to me a natural thing to include next.
This work does a good job of separating risk scoring from continuous monitoring and shows why looking at the full sequence of agent actions can provide more information than checking individual events. I also liked that the paper is careful about the early-warning results. The main limitation is that the current testing uses constructed scenarios and example thresholds. Testing with larger real-world datasets, realistic false-positive rates, and actual response times would help show whether this approach performs better in practice.
Cite this project
@misc{li2026towards,
title = {{Towards Early Warning for AI Agent Incidents: Monitoring the Dynamics of Risk Trajectories}},
author = {Yifan Li and John Gia Bao Luc and Rashik Shahjahan and Kai Ting},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/towards-early-warning-for-ai-agent-incidents-monitoring-the-dynamics-of-risk-trajectories-lxdt}},
url = {https://apartresearch.com/sprints/projects/towards-early-warning-for-ai-agent-incidents-monitoring-the-dynamics-of-risk-trajectories-lxdt}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …