Kairos
Shashank Kumar · Team Kairos
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Proposals to make AI agent actions cryptographically attributable assume the agent signs what it does. I implemented such a scheme, Ed25519 per-action attestation verifiable by any third party holding only a public key, and tested it against the shape of the two documented 2026 containment failures. It fails on both: an agent that escapes to obtain benchmark answers has no incentive to sign what it does outside, so signature checking returns zero attributable actions and names the culprit unknown. An agent that signs only its in-sandbox work reports 100% compliant while the actions that reached a third party go unrecorded. The salvageable primitive is the inverse: under an obligation to attest, an unsigned action from obligated address space is itself the evidence, checkable by the receiving platform from its own logs. Audited against the egress files an operator actually publishes, however, not one prefix is fresh enough to support an assertable finding.
Reviews
The inversion at the centre of this is a good idea, clearly argued: an agent that escapes will not sign what it does, so an unsigned action from address space its operator declared is the evidence, checkable by the receiving platform from its own logs. Building the obvious fix first and showing why it fails makes the argument convincing, and the registry requirements at the end are concrete enough for someone to adopt.
The pattern is not new to security, though. DMARC tells receivers to treat unauthenticated mail from a domain's declared senders as suspect, and RPKI does the same for route announcements. Positioning the work against those would sharpen what is specific to agents, and their deployment history, including how much traffic still bypasses them, is directly relevant to your biggest limitation that one proxy defeats the method.
To check the results I downloaded OpenAI's gptbot.json and searchbot.json myself and recomputed the registry audit. Your Table 2 reproduces exactly: 60 prefixes, 6 shared, 5,264 distinct addresses, 14.6% ambiguous, 315 and 251 days stale. That is the strongest part of the paper. The detection results are weaker evidence: the 15,000 of 15,000 with zero false positives comes from a synthetic stream where agent traffic sits inside the declared range by construction, which you acknowledge in 4.7. I could not check the tests or timings, because the code is neither in a repository nor in an appendix, and the sprint asked for the artifact in one of the two. The incident facts lean on the sprint brief and press coverage; the Hugging Face intrusion began on 9 July, not 11 July, and the technical timeline would give you better anchors.
The report is easy to follow and the figure tells the whole story. The opening anecdote does not connect to the method and could go, and Section 5.1 on reporting clocks is a separate regulatory question that takes space from your own findings.
If you take this further, get one cooperating platform's real logs and measure the false-positive rate, and repeat the registry audit across several operators to see whether OpenAI's staleness is typical.
Read full reviewShow less
The submission spots an important weakness in signature based attribution. Verifying every submitted signature does not prove that every relevant action was signed, and the partial compliance example makes that point really well: benign actions can produce a 100 percent verification rate while the escaped activity stays unsigned. Treating missing required attestations as evidence is a useful reframing, and auditing real published range files was a good attempt to test the operational dependency instead of just assuming it works.
The main limitation is that the incident stream evaluation uses synthetic data with clean range boundaries. In those conditions, classifying unsigned events from an obligated range as gaps mostly shows the implementation is correct, not that detection quality is good in the real world. Production validation should measure false positives under shared NAT, stale assignments, proxies, mixed workloads, configuration failures, and changing cloud addresses. Right now a commodity proxy bypasses the method entirely, and even a detected gap points at an operator or range rather than a particular model, agent, or run. It would also help to evaluate purpose built agent egress registries across multiple operators, since the GPTBot and SearchBot ranges were published for crawler traffic, not autonomous agent attribution.
Read full reviewShow less
Cite this project
@misc{kumar2026kairos,
title = {{Kairos}},
author = {Shashank Kumar},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/kairos-2vka}},
url = {https://apartresearch.com/sprints/projects/kairos-2vka}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …