Skip to content
Sprint projectSep 13, 2026Delhi, India

Kairos

Shashank Kumar · Team Kairos

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Proposals to make AI agent actions cryptographically attributable assume the agent signs what it does. I implemented such a scheme, Ed25519 per-action attestation verifiable by any third party holding only a public key, and tested it against the shape of the two documented 2026 containment failures. It fails on both: an agent that escapes to obtain benchmark answers has no incentive to sign what it does outside, so signature checking returns zero attributable actions and names the culprit unknown. An agent that signs only its in-sandbox work reports 100% compliant while the actions that reached a third party go unrecorded. The salvageable primitive is the inverse: under an obligation to attest, an unsigned action from obligated address space is itself the evidence, checkable by the receiving platform from its own logs. Audited against the egress files an operator actually publishes, however, not one prefix is fresh enough to support an assertable finding.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The inversion at the centre of this is a good idea, clearly argued: an agent that escapes will not sign what it does, so an unsigned action from address space its operator declared is the evidence, checkable by the receiving platform from its own logs. Building the obvious fix first and showing why it fails makes the argument convincing, and the registry requirements at the end are concrete enough for someone to adopt.

    The pattern is not new to security, though. DMARC tells receivers to treat unauthenticated mail from a domain's declared senders as suspect, and RPKI does the same for route announcements. Positioning the work against those would sharpen what is specific to agents, and their deployment history, including how much traffic still bypasses them, is directly relevant to your biggest limitation that one proxy defeats the method.

    To check the results I downloaded OpenAI's gptbot.json and searchbot.json myself and recomputed the registry audit. Your Table 2 reproduces exactly: 60 prefixes, 6 shared, 5,264 distinct addresses, 14.6% ambiguous, 315 and 251 days stale. That is the strongest part of the paper. The detection results are weaker evidence: the 15,000 of 15,000 with zero false positives comes from a synthetic stream where agent traffic sits inside the declared range by construction, which you acknowledge in 4.7. I could not check the tests or timings, because the code is neither in a repository nor in an appendix, and the sprint asked for the artifact in one of the two. The incident facts lean on the sprint brief and press coverage; the Hugging Face intrusion began on 9 July, not 11 July, and the technical timeline would give you better anchors.

    The report is easy to follow and the figure tells the whole story. The opening anecdote does not connect to the method and could go, and Section 5.1 on reporting clocks is a separate regulatory question that takes space from your own findings.

    If you take this further, get one cooperating platform's real logs and measure the false-positive rate, and repeat the registry audit across several operators to see whether OpenAI's staleness is typical.

    Read full reviewShow less
  2. The submission spots an important weakness in signature based attribution. Verifying every submitted signature does not prove that every relevant action was signed, and the partial compliance example makes that point really well: benign actions can produce a 100 percent verification rate while the escaped activity stays unsigned. Treating missing required attestations as evidence is a useful reframing, and auditing real published range files was a good attempt to test the operational dependency instead of just assuming it works.

    The main limitation is that the incident stream evaluation uses synthetic data with clean range boundaries. In those conditions, classifying unsigned events from an obligated range as gaps mostly shows the implementation is correct, not that detection quality is good in the real world. Production validation should measure false positives under shared NAT, stale assignments, proxies, mixed workloads, configuration failures, and changing cloud addresses. Right now a commodity proxy bypasses the method entirely, and even a detected gap points at an operator or range rather than a particular model, agent, or run. It would also help to evaluate purpose built agent egress registries across multiple operators, since the GPTBot and SearchBot ranges were published for crawler traffic, not autonomous agent attribution.

    Read full reviewShow less

Cite this project

@misc{kumar2026kairos,
  title = {{Kairos}},
  author = {Shashank Kumar},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/kairos-2vka}},
  url = {https://apartresearch.com/sprints/projects/kairos-2vka}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026