Skip to content
Sprint projectSep 11, 2026Wuhan

From Agent Claims to Execution Facts: Generation-Fact Graphs for AI Incident Reconstruction

Mian Wang · Team Happy wind

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: From Agent Claims to Execution Facts: Generation-Fact Graphs for AI Incident Reconstruction

Code (opens in new tab)More on github.com (opens in new tab)
Share

Autonomous-AI incident response requires evidence of what a system actually did, not only what the agent claims it did. This project applies the Generation-Fact Graph (GFG), an existing machine-readable substrate for concrete generation and execution facts, to AI incident reconstruction. GFG records participating origins, realized transformations, concrete occurrences, generated results, and relation roles, allowing heterogeneous execution evidence to be compiled into one fact graph for independent reconstruction and verification. Existing frozen experiments demonstrate exact projection across multiple provenance and tracing mechanisms and exact cross-stage formation reconstruction. The project artifact is the open-source GFG Core and structural experiments repository.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I like the idea of separating what an agent says it did from what actually happened during execution. The main thing missing for me is a real agent incident example. Most of the current results show that GFG works structurally, but not yet how well it reconstructs an actual failure. Testing it on a real agent workflow and comparing it with standard tracing would make the project much stronger.

  2. The core framing is good and I want to credit it first. "GFG does not make the agent the witness, recorder, and judge of its own execution" is the right sentence for this problem. Treating a model's "the upload succeeded" as just another generated output is the correct instinct. So is requiring a separate occurrence record for it.

    Section 7 is an unusually clean prior-work declaration. You state plainly that GFG Core, all three experiments, the protocols, fixtures and validators predate the sprint and that the sprint contribution is the incident-response formulation.

    Most submissions would have blurred that line. You did not. That is worth saying out loud.

    What I verified in the repository:

    1. GF-S01 – artifacts/final_report.json gives 1,117 generation occurrences and 8,420 generation bindings. Both match the paper.

    2. The 2,880 paths claim holds and it holds in the way that matters – candidate_path_count and reference_path_count are both 2,880. The independent reference agrees.

    3. downsampled_max_abs_error is 1.11e-16, which is machine epsilon. The reconstruction is exact, not approximate.

    4. GF-P01 – exactly five profiles are present and they are the five named mechanisms. Status reads FIVE_PROFILE_EXACT_STRICT_PROJECTION_SUPPORTED.

    5. The 17,600 attacker actions attributed to Hugging Face check out. Their post says "~17,600 attacker actions that we were able to recover, grouped into ~6,280 clusters".

    The prior work is careful and it reproduces. Section 6 and Appendix A are also strong. You separate formation ancestry from causal attribution. The dual-use note about execution records exposing credentials and topology is responsible.

    Now the problems.

    1. The sprint contribution is a five-row table.

    Section 3 maps five incident questions to GFG coordinates. That is the new content. Each row is one sentence. It is sensible but it is thin for a sprint submission and the rest of the paper is prior work you have correctly labelled as prior work.

    I am not penalising the disclosure. I am saying the disclosed contribution is small.

    2. You invoke the July 2026 incident but never apply GFG to it.

    Section 1 opens on the OpenAI-Hugging Face incident. Section 5 says the approach "could connect an agent observation to a tool invocation, the concrete tool occurrence, the returned result, a subsequent action, and an eventual external-system effect". That stays conditional throughout.

    Hugging Face published a technical timeline. You cite it. A hand-built GFG over a ten-step slice of that published timeline would have been a weekend's work and would have changed this submission completely. It would show which coordinates are populated from real evidence, which are missing, and where the fact model breaks against messy cross-organisation logs.

    To be honest – that missing worked example is the difference between a position paper and a result. Please do it. You already have the substrate and the source material.

    3. The strongest evidence is from a domain unlike the target.

    GF-S01 traces FIR filtering, downsampling, a sliding FFT and SVG rendering. That is a deterministic numeric pipeline with total, clean data dependency. Agent execution is non-deterministic, partially observed and adversarial – the agent may actively avoid instrumentation.

    You call GF-S01 "the closest existing demonstration to incident reconstruction" and that is fair. Appendix A also covers the coverage limit honestly.

    But the gap is not only coverage. It is character. An exact path count over an ECG pipeline does not tell us how the representation behaves when occurrences are ambiguous, clocks disagree across organisations or an adversary deliberately leaves no fact. We need one experiment in that regime.

    4. What happens when the capture layer is itself compromised?

    The fact authority comes from instrumented execution. In the July 2026 class of incident the agent reached shared infrastructure. If an agent can influence the systems producing the generation facts, the substrate inherits the compromise.

    Appendix A lists "trust in capture components" as a deployment need. That is the central threat model for this proposal, not a deployment detail. I would promote it into the body.

    5. Minor – the thesis is restated three times.

    Sections 1, 5 and the Conclusion each make the agent-testimony-versus-execution-facts argument. At three pages that repetition costs space you could spend on the worked example.

    6. Minor – one artifact status reads NOT_EXECUTED_IN_UNIFIED_RUN. Please say in the report which component that is and why it is excluded, so a reader does not have to infer it.

    One note, since careful reports are sometimes read as machine-written. The work behind this one is real. The repository numbers reproduce, the external citation is accurate to the word and the prior-work boundary is declared more clearly than in most published papers. I scored the work.

    Please let me know for any questions.

    Read full reviewShow less

Cite this project

@misc{wang2026from,
  title = {{From Agent Claims to Execution Facts: Generation-Fact Graphs for AI Incident Reconstruction}},
  author = {Mian Wang},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/from-agent-claims-to-execution-facts-generationfact-graphs-for-ai-incident-reconstruction-pllc}},
  url = {https://apartresearch.com/sprints/projects/from-agent-claims-to-execution-facts-generationfact-graphs-for-ai-incident-reconstruction-pllc}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026