Breadcrumbs
Seth Ferreira · Team breadcrumbs
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Breadcrumbs is an innovative tool that allows monitoring all the decisions and actions of an AI agent and classifying them as log-derived, guessed, or synthetic trials. It helps track what went wrong and understand why a particular play was selected. A simple rule-based classifier was tested on a standard set of 1000 trajectories of the agent’s behavior. Initially, only the actions were taken into account, and then – the resources and tools used. As a result, it was found that the more context was given to the classifier, the higher its accuracy was in identifying incidents. However, the true positives were competing with the false ones due to the high number of various tools used by the agent, even when it did nothing wrong. This was found to be an inherent property of such classifiers, which cannot differentiate between new and harmful actions. This discovery might be helpful to other researchers in the field of AI safety who are developing better classifiers.
Reviews
Breadcrumbs reports a negative result honestly and backs it with a reproducible pipeline. The schema, importer, detector and threshold sweep are all in the repository. Separating observed from inferred fields is a sensible foundation for incident forensics, and the section on what did not work is worth reading.
The problem is real, but flagging whatever is new within a trajectory is a well-established approach in log anomaly detection, and the report does not connect the work to the July 2026 incident this sprint is about. One paragraph mapping the schema onto the agent actions in the Hugging Face technical timeline would show what Breadcrumbs adds for a responder. A Related Work section with actual citations would show what is new.
To check the results I cloned the repository, read the detector and the ATBench importer, and compared the committed results files against the tables in the report. The threshold sweep in results/threshold_sensitivity.csv matches Table 2 exactly. The central finding, however, may be an artifact of the importer. atbench_loader.py stores the full JSON arguments of each tool call as the resource, so almost every call counts as a new resource and the detector flags nearly every trajectory. Rerunning with a coarser resource, such as tool name plus target, would test whether the tripwire effect survives. Table 1 reports precision 0.99 and FPR 0.00 for configuration A, while results/ablation_table.csv gives 0.68 and 0.107. With the default weights and threshold 3 in detector.py, configuration A can only fire on the escalation keywords, which is worth stating. The threshold sweep keeps the weights fixed, so it cannot answer the untuned-weights limitation. I read the code without running it.
The report is easy to follow, with problem, method, findings and limitations clearly in place. Before sharing it further, remove the citation placeholder in Related Work and fix the Figure 1 caption, since the plot shows precision near 0.50 through threshold 5. On page 5, "increased precision" should read recall. List the detector's five features explicitly; the text currently says "including at least". The sprint asked for a Limitations and Dual-Use appendix; the limitations are in the main text, but the dual-use considerations are missing.
If you take this further, define expected tool and resource usage against an outside reference, such as the dataset-wide distribution, and test it on a second benchmark.
Read full reviewShow less
Negative result is clear (novelty boosts recall but flags everything), but project stops a bit early limiting contribution. For example, could the failure mode be fixed? For example, replace "new within the trajectory" with an external baseline of what is expected (allowed tools for the task, patterns from true-safe runs"), then see if the detector improves.
Good direction.
Cite this project
@misc{ferreira2026breadcrumbs,
title = {{Breadcrumbs}},
author = {Seth Ferreira},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/breadcrumbs-z0rv}},
url = {https://apartresearch.com/sprints/projects/breadcrumbs-z0rv}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …