Skip to content
Sprint projectSep 14, 2026Mexico (remote)

Scope before chronology

Roman Vinogradov · Team Scope before chronology

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

An offline checker for comparing AI incident claims without silently dropping event scope or evidence status. The artifact includes eight source-linked encodings, four illustrative comparisons, 15 passing unit tests and 28 authored controls. These demonstrate software behavior, not independent real-world accuracy. This is an explicitly AI-led solo entry under Roman Vinogradov, the human operator and prize recipient. Codex performed research, coding, testing and report drafting; no independent human review is claimed.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is a small, honest artifact with unusually precise limitations. The checker is clean, dependency-free and does exactly what the report says, and I confirmed the Hugging Face encodings against the timeline: 13:37 last meaningful activity, 14:14 last logged event, node root at 19:53 on day 3.

    The problem is real, since cross-source briefs do flatten different kinds of timestamps, but the contribution is a type check before comparison, which is a well-established pattern. The value would show only in use by analysts on their own briefs.

    To check the results I cloned the repository, read claimcheck.py, the tests and all eight encoded claims, and compared the saved outputs with Table 1. They match. The comparison itself is decided by the encoding: each pair was given different predicate or scope labels by the same agent that selected the pairs, so not_comparable follows by construction, and the interval-only baseline flags differences no analyst would call conflicts. The independent-annotation study you describe is the test that would show whether the checker catches anything a careful reader would miss.

    The report is short, clear and well structured, with exactly the detail needed.

    If you take this further, have two analysts encode the same passages independently and see whether the checker surfaces disagreements they did not notice.

    Read full reviewShow less
  2. I like the problem you are pointing at here. It is very easy for incident summaries to turn two timestamps into an apparent contradiction when the underlying sources are actually talking about different events, systems, scopes, or evidence states. The paper makes that failure mode easy to understand.

    The prototype also seems appropriately modest. On your four examples, forcing claims to carry system, subject, predicate, scope, and evidence metadata changes three apparent timestamp conflicts into "not comparable" and sends the inferred-versus-reported case to review. That is exactly the kind of behavior I would want from a tool like this: abstention rather than confidently manufacturing a contradiction.

    My main question is whether this needs to be software rather than a disciplined structured-analysis template. The underlying comparison logic is intentionally simple, and I would have liked a baseline against something like an ordinary spreadsheet or structured template used by an experienced incident analyst. If the tool reduces analyst errors or makes review materially faster, that would strengthen the Impact Potential & Innovation case a lot.

    The other major limitation is that the same agent produced the claim encodings, implementation, controls, and expected outputs. As you acknowledge, the biggest source of error may be the encoding of the source claim rather than the interval-comparison code itself.

    Very clear paper overall! I appreciated how little it overclaims. :)

    Read full reviewShow less

Cite this project

@misc{vinogradov2026scope,
  title = {{Scope before chronology}},
  author = {Roman Vinogradov},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/scope-before-chronology-2fd2}},
  url = {https://apartresearch.com/sprints/projects/scope-before-chronology-2fd2}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026