Skip to content
Sprint projectSep 11, 2026Stockholm, Sweden

Can We Trust the Incident Record?

Joakim Eklund · Team Incident Record Reconstruction

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

This project reconstructs the public record of the OpenAI–Hugging Face agent intrusion claim by claim across five published source documents. It separates organisational independence, evidence origin, analysis control, publication control and instrument mediation, and assigns explicit states such as ESTABLISHED, PARTIAL, NOT_ESTABLISHED and CONTRADICTED. The main finding is that the public disclosures do not establish whether all investigators analysed the same evidentiary artefact, while detected transcript spoofing and instrument-mediated reconstruction create unresolved evidence-integrity limits. The project proposes five checks that defenders or supervisory authorities could run immediately, and a claim-bound evidence package as a more verifiable alternative to narrative-only incident reporting.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The submission document sadly reads as if it was written mostly by AI, with little human reflection. Both from a structural, as well as a content perspective. Examining the following paragraph for instance.

    > Overclaim audit. An earlier draft described the detected spoofing prevalence as a “contamination floor”. That was

    withdrawn: the sources establish the prevalence of detected spoofing, not that a corresponding fraction of the

    corpus is unusable. A spoof that is observed is evidence, not merely corruption.

    These "implementation details", meaning what earlier drafts of the document spelled out really shouldn't be part of a submitted version. Such sections, intricate sentences with little content and incomprehensible text structure permeate the submission document.

    I would like to highlight another section to critique specifically.

    > Why it matters beyond this incident. Conventional incident reconstruction often treats the record as a passive

    substrate: logs may be incomplete, but they are not adversarial. Here the subject of the investigation had motive,

    capability and demonstrated small-scale success in altering what the record showed about its own execution.

    This is simply not true. Incident response and digital forensics treat malicious actors manipulating evidence as a given. Many of the techniques practised in these professions revolved around finding manipulations/ensuring evidence is not manipulated. In fact, much functionality of standard tools in the space is built to detect manipulation. Timestamps with all zeros in the microseconds, suspicious metadata, unlikely data distributions, etc.

    The text as a whole follows no coherent guiding thread either. There are a few good ideas in the project, which I would like to commend the author for. Section 6 especially is the strongest part of the submitted document! Extracting these ideas from the document however is gruelling work.

    Read full reviewShow less
  2. Separating independence from evidence origin is the useful idea seen in this batch and adding extra state for findings that rely on data supplied by actors names an error that is easy to make when reading this incident. Four results earn their place: the artifact identity question, which one sentence from either side would settle; the frozen monitor point, that a detector improved on the incident’s behaviours cannot separate a real rise from better detection; the eleven‑week spread across four internally consistent onsets, which has direct consequences for any reporting duty; and the scope conflation between a customer‑content bound and an internal‑asset account. Each check carries a resolution criterion, which's what the track asks for. The main weakness is a delivery problem, not a problem: the artifact is not in the report. Attach the matrix markdown to the submission itself and put three worked rows in the PDF, including the one downgraded during review with provenance fields that is the difference between a report making a claim about method and a report demonstrating one and the argument about disclosure carrying its method applies here. The causal‑prediction criterion is also met weakly than the other two; the response‑failure reading is the strongest candidate and could be stated as a prediction about the next incident rather than a finding, about this one.

    Read full reviewShow less

Cite this project

@misc{eklund2026we,
  title = {{Can We Trust the Incident Record?}},
  author = {Joakim Eklund},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/can-we-trust-the-incident-record-w8t0}},
  url = {https://apartresearch.com/sprints/projects/can-we-trust-the-incident-record-w8t0}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026