From Agent Claims to Execution Facts: Generation-Fact Graphs for AI Incident Reconstruction
Mian Wang · Team Happy wind
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Autonomous-AI incident response requires evidence of what a system actually did, not only what the agent claims it did. This project applies the Generation-Fact Graph (GFG), an existing machine-readable substrate for concrete generation and execution facts, to AI incident reconstruction. GFG records participating origins, realized transformations, concrete occurrences, generated results, and relation roles, allowing heterogeneous execution evidence to be compiled into one fact graph for independent reconstruction and verification. Existing frozen experiments demonstrate exact projection across multiple provenance and tracing mechanisms and exact cross-stage formation reconstruction. The project artifact is the open-source GFG Core and structural experiments repository.
Reviews
I like the idea of separating what an agent says it did from what actually happened during execution. The main thing missing for me is a real agent incident example. Most of the current results show that GFG works structurally, but not yet how well it reconstructs an actual failure. Testing it on a real agent workflow and comparing it with standard tracing would make the project much stronger.
The core framing is good and I want to credit it first. "GFG does not make the agent the witness, recorder, and judge of its own execution" is the right sentence for this problem. Treating a model's "the upload succeeded" as just another generated output is the correct instinct. So is requiring a separate occurrence record for it.
Section 7 is an unusually clean prior-work declaration. You state plainly that GFG Core, all three experiments, the protocols, fixtures and validators predate the sprint and that the sprint contribution is the incident-response formulation.
Most submissions would have blurred that line. You did not. That is worth saying out loud.
What I verified in the repository:
1. GF-S01 – artifacts/final_report.json gives 1,117 generation occurrences and 8,420 generation bindings. Both match the paper.
2. The 2,880 paths claim holds and it holds in the way that matters – candidate_path_count and reference_path_count are both 2,880. The independent reference agrees.
3. downsampled_max_abs_error is 1.11e-16, which is machine epsilon. The reconstruction is exact, not approximate.
4. GF-P01 – exactly five profiles are present and they are the five named mechanisms. Status reads FIVE_PROFILE_EXACT_STRICT_PROJECTION_SUPPORTED.
5. The 17,600 attacker actions attributed to Hugging Face check out. Their post says "~17,600 attacker actions that we were able to recover, grouped into ~6,280 clusters".
The prior work is careful and it reproduces. Section 6 and Appendix A are also strong. You separate formation ancestry from causal attribution. The dual-use note about execution records exposing credentials and topology is responsible.
Now the problems.
1. The sprint contribution is a five-row table.
Section 3 maps five incident questions to GFG coordinates. That is the new content. Each row is one sentence. It is sensible but it is thin for a sprint submission and the rest of the paper is prior work you have correctly labelled as prior work.
I am not penalising the disclosure. I am saying the disclosed contribution is small.
2. You invoke the July 2026 incident but never apply GFG to it.
Section 1 opens on the OpenAI-Hugging Face incident. Section 5 says the approach "could connect an agent observation to a tool invocation, the concrete tool occurrence, the returned result, a subsequent action, and an eventual external-system effect". That stays conditional throughout.
Hugging Face published a technical timeline. You cite it. A hand-built GFG over a ten-step slice of that published timeline would have been a weekend's work and would have changed this submission completely. It would show which coordinates are populated from real evidence, which are missing, and where the fact model breaks against messy cross-organisation logs.
To be honest – that missing worked example is the difference between a position paper and a result. Please do it. You already have the substrate and the source material.
3. The strongest evidence is from a domain unlike the target.
GF-S01 traces FIR filtering, downsampling, a sliding FFT and SVG rendering. That is a deterministic numeric pipeline with total, clean data dependency. Agent execution is non-deterministic, partially observed and adversarial – the agent may actively avoid instrumentation.
You call GF-S01 "the closest existing demonstration to incident reconstruction" and that is fair. Appendix A also covers the coverage limit honestly.
But the gap is not only coverage. It is character. An exact path count over an ECG pipeline does not tell us how the representation behaves when occurrences are ambiguous, clocks disagree across organisations or an adversary deliberately leaves no fact. We need one experiment in that regime.
4. What happens when the capture layer is itself compromised?
The fact authority comes from instrumented execution. In the July 2026 class of incident the agent reached shared infrastructure. If an agent can influence the systems producing the generation facts, the substrate inherits the compromise.
Appendix A lists "trust in capture components" as a deployment need. That is the central threat model for this proposal, not a deployment detail. I would promote it into the body.
5. Minor – the thesis is restated three times.
Sections 1, 5 and the Conclusion each make the agent-testimony-versus-execution-facts argument. At three pages that repetition costs space you could spend on the worked example.
6. Minor – one artifact status reads NOT_EXECUTED_IN_UNIFIED_RUN. Please say in the report which component that is and why it is excluded, so a reader does not have to infer it.
One note, since careful reports are sometimes read as machine-written. The work behind this one is real. The repository numbers reproduce, the external citation is accurate to the word and the prior-work boundary is declared more clearly than in most published papers. I scored the work.
Please let me know for any questions.
Read full reviewShow less
Cite this project
@misc{wang2026from,
title = {{From Agent Claims to Execution Facts: Generation-Fact Graphs for AI Incident Reconstruction}},
author = {Mian Wang},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/from-agent-claims-to-execution-facts-generationfact-graphs-for-ai-incident-reconstruction-pllc}},
url = {https://apartresearch.com/sprints/projects/from-agent-claims-to-execution-facts-generationfact-graphs-for-ai-incident-reconstruction-pllc}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …