Skip to content
Sprint projectSep 13, 2026CAPE TOWN

Breadcrumbs

Seth Ferreira · Team breadcrumbs

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Breadcrumbs is an innovative tool that allows monitoring all the decisions and actions of an AI agent and classifying them as log-derived, guessed, or synthetic trials. It helps track what went wrong and understand why a particular play was selected. A simple rule-based classifier was tested on a standard set of 1000 trajectories of the agent’s behavior. Initially, only the actions were taken into account, and then – the resources and tools used. As a result, it was found that the more context was given to the classifier, the higher its accuracy was in identifying incidents. However, the true positives were competing with the false ones due to the high number of various tools used by the agent, even when it did nothing wrong. This was found to be an inherent property of such classifiers, which cannot differentiate between new and harmful actions. This discovery might be helpful to other researchers in the field of AI safety who are developing better classifiers.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Breadcrumbs reports a negative result honestly and backs it with a reproducible pipeline. The schema, importer, detector and threshold sweep are all in the repository. Separating observed from inferred fields is a sensible foundation for incident forensics, and the section on what did not work is worth reading.

    The problem is real, but flagging whatever is new within a trajectory is a well-established approach in log anomaly detection, and the report does not connect the work to the July 2026 incident this sprint is about. One paragraph mapping the schema onto the agent actions in the Hugging Face technical timeline would show what Breadcrumbs adds for a responder. A Related Work section with actual citations would show what is new.

    To check the results I cloned the repository, read the detector and the ATBench importer, and compared the committed results files against the tables in the report. The threshold sweep in results/threshold_sensitivity.csv matches Table 2 exactly. The central finding, however, may be an artifact of the importer. atbench_loader.py stores the full JSON arguments of each tool call as the resource, so almost every call counts as a new resource and the detector flags nearly every trajectory. Rerunning with a coarser resource, such as tool name plus target, would test whether the tripwire effect survives. Table 1 reports precision 0.99 and FPR 0.00 for configuration A, while results/ablation_table.csv gives 0.68 and 0.107. With the default weights and threshold 3 in detector.py, configuration A can only fire on the escalation keywords, which is worth stating. The threshold sweep keeps the weights fixed, so it cannot answer the untuned-weights limitation. I read the code without running it.

    The report is easy to follow, with problem, method, findings and limitations clearly in place. Before sharing it further, remove the citation placeholder in Related Work and fix the Figure 1 caption, since the plot shows precision near 0.50 through threshold 5. On page 5, "increased precision" should read recall. List the detector's five features explicitly; the text currently says "including at least". The sprint asked for a Limitations and Dual-Use appendix; the limitations are in the main text, but the dual-use considerations are missing.

    If you take this further, define expected tool and resource usage against an outside reference, such as the dataset-wide distribution, and test it on a second benchmark.

    Read full reviewShow less
  2. Negative result is clear (novelty boosts recall but flags everything), but project stops a bit early limiting contribution. For example, could the failure mode be fixed? For example, replace "new within the trajectory" with an external baseline of what is expected (allowed tools for the task, patterns from true-safe runs"), then see if the detector improves.

    Good direction.

Cite this project

@misc{ferreira2026breadcrumbs,
  title = {{Breadcrumbs}},
  author = {Seth Ferreira},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/breadcrumbs-z0rv}},
  url = {https://apartresearch.com/sprints/projects/breadcrumbs-z0rv}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026