Scope before chronology
Roman Vinogradov · Team Scope before chronology
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
An offline checker for comparing AI incident claims without silently dropping event scope or evidence status. The artifact includes eight source-linked encodings, four illustrative comparisons, 15 passing unit tests and 28 authored controls. These demonstrate software behavior, not independent real-world accuracy. This is an explicitly AI-led solo entry under Roman Vinogradov, the human operator and prize recipient. Codex performed research, coding, testing and report drafting; no independent human review is claimed.
Reviews
This is a small, honest artifact with unusually precise limitations. The checker is clean, dependency-free and does exactly what the report says, and I confirmed the Hugging Face encodings against the timeline: 13:37 last meaningful activity, 14:14 last logged event, node root at 19:53 on day 3.
The problem is real, since cross-source briefs do flatten different kinds of timestamps, but the contribution is a type check before comparison, which is a well-established pattern. The value would show only in use by analysts on their own briefs.
To check the results I cloned the repository, read claimcheck.py, the tests and all eight encoded claims, and compared the saved outputs with Table 1. They match. The comparison itself is decided by the encoding: each pair was given different predicate or scope labels by the same agent that selected the pairs, so not_comparable follows by construction, and the interval-only baseline flags differences no analyst would call conflicts. The independent-annotation study you describe is the test that would show whether the checker catches anything a careful reader would miss.
The report is short, clear and well structured, with exactly the detail needed.
If you take this further, have two analysts encode the same passages independently and see whether the checker surfaces disagreements they did not notice.
Read full reviewShow less
I like the problem you are pointing at here. It is very easy for incident summaries to turn two timestamps into an apparent contradiction when the underlying sources are actually talking about different events, systems, scopes, or evidence states. The paper makes that failure mode easy to understand.
The prototype also seems appropriately modest. On your four examples, forcing claims to carry system, subject, predicate, scope, and evidence metadata changes three apparent timestamp conflicts into "not comparable" and sends the inferred-versus-reported case to review. That is exactly the kind of behavior I would want from a tool like this: abstention rather than confidently manufacturing a contradiction.
My main question is whether this needs to be software rather than a disciplined structured-analysis template. The underlying comparison logic is intentionally simple, and I would have liked a baseline against something like an ordinary spreadsheet or structured template used by an experienced incident analyst. If the tool reduces analyst errors or makes review materially faster, that would strengthen the Impact Potential & Innovation case a lot.
The other major limitation is that the same agent produced the claim encodings, implementation, controls, and expected outputs. As you acknowledge, the biggest source of error may be the encoding of the source claim rather than the interval-comparison code itself.
Very clear paper overall! I appreciated how little it overclaims. :)
Read full reviewShow less
Cite this project
@misc{vinogradov2026scope,
title = {{Scope before chronology}},
author = {Roman Vinogradov},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/scope-before-chronology-2fd2}},
url = {https://apartresearch.com/sprints/projects/scope-before-chronology-2fd2}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …