ScopeAI
Tammy
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This project investigates: How completely do the public investigations cover a predefined set of incident-accountability questions, who set their boundaries, and which answers remain dependent on the operator's own account?
Reviews
This research is useful to highlight the gaps exposed by the lack of legal requirements for comprehensive third party access to labs' internal work. More policymakers need to be aware of these scope limitations as it pertains to "after action reports." Would love to see this frameworks opened up and generalized a bit such that it is applicable to future incidents without having to re-write the questions. I really like the idea of a Disclosure Minimum.
This work makes a useful distinction between independent analysis and independently determined investigation scope, and turns it into a practical disclosure checklist.
Its treatment of the external investigators’ contributions and limitations is balanced. Clearer criteria for what counts as full coverage, including the distinction between remediation plans and demonstrated effectiveness, would make the coding easier to interpret and reproduce. An independent coding pass and a small sensitivity analysis would strengthen confidence in the headline percentages. Larger text and more selective use of tables would also improve readability.
I think ScopeAI makes a useful distinction because an independent investigation can still leave important questions outside its assigned scope. The paper maps which questions an external AI incident assessment addresses and proposes requests for further information. I appreciated that it recognizes the assessment’s value rather than treating unanswered questions as evidence that the investigators lacked independence.
The numerical findings need correction before others rely on them. The scores shown in the paper do not match one of its reported percentages, which makes the size of the claimed gap uncertain. I would reconcile those figures and ask another reviewer to apply the scoring rules to the same sources. The broader concern is worth examining, but missing coverage alone does not establish misconduct or unreliable findings.
Cite this project
@misc{tammy2026scopeai,
title = {{ScopeAI}},
author = {Tammy},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/scopeai-v1hd}},
url = {https://apartresearch.com/sprints/projects/scopeai-v1hd}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …