Detection Was Not the Failure: A Replayable Conformance Harness for the July 2026 Agent Intrusion1
Chijioke Ubajaka
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Hugging Face's forensic reconstruction of the July 2026 agent intrusion records that its security stack correlated the attack into a coherent signal and then failed to raise criticality and page the on-call team. I treat that as the primary failure and test it. From published aggregates: per-phase action counts, daily totals and seventeen timestamped anchors, I reconstruct a defender-visible event stream by iterative proportional fitting, generate 30 days of synthetic benign platform workload, and score four escalation policies on detection point and false-page cost. A per-event severity policy never fires. A boundary-accumulation policy pages 51.5 hours before containment but costs 69 false pages per week, driven by the deploy pipeline. A principal-scoped variant retains that lead at 0.5 false pages per week, but only while the principal-to-boundary map stays accurate to better than one violation in ten thousand sessions. The bottleneck is identity hygiene, not detection logic.

Reviews
This is a strong submission, and unusually reproducible. Turning a real incident into a testable detection experiment is genuinely hard to pull off, and the comparison between detection lead time and the cost of false pages is the most valuable part. The unreconciled event counts are handled up front instead of glossed over, and the pinned timestamps, the mix of benign and attack traffic, the sensitivity testing, regression guards, and the honest limitations section all make the results easier to trust.
The main gap is external validity. The benign workload, defender visibility, session linkage, and map drift rate are all synthetic or assumed, and the principal scoped policy only works if the identity map stays extremely accurate, which is a big ask in practice. A good next step would be validating against even a small sample of real operational telemetry, and spelling out how event severity and rate baselines were assigned. It would also be worth testing identity rotation, fragmented sessions, and activity timed around legitimate deployment windows, to see whether the approach holds up against an adaptive agent.
Read full reviewShow less
This study gives containment discussions a useful testbed by comparing warning time with the cost of false alerts. Including legitimate deployments that cross trust boundaries makes the comparison more realistic, and the report explains its assumptions clearly. The findings are specific to the simulated workload, with the severity choices and hourly alert timing limiting the policy comparison. The released harness makes those assumptions inspectable and gives the work value beyond its headline numbers.
Cite this project
@misc{ubajaka2026detection,
title = {{Detection Was Not the Failure: A Replayable Conformance Harness for the July 2026 Agent Intrusion1}},
author = {Chijioke Ubajaka},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detection-was-not-the-failure-a-replayable-conformance-harness-for-the-july-2026-agent-intrusion1-pp2o}},
url = {https://apartresearch.com/sprints/projects/detection-was-not-the-failure-a-replayable-conformance-harness-for-the-july-2026-agent-intrusion1-pp2o}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …