JAIL: Justified Artifact Investigation Layer, A Provenance-Bound Authorization Layer for AI Incident-Response Agents
Aref khalil · Team Checkpoint
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
AI incident-response agents routinely encounter the exact artifacts they're trained to refuse: malware, exploit code, credential-theft scripts, attacker-authored content. AIL separates four properties that trusted user exceptions usually collapse into one: identity (who's asking), provenance (where the artifact came from), integrity (is it the same bytes that were recorded), and authorization scope (what they're allowed to ask for). An authorization token is only issued when all four check out together, and the token can only ever unlock analysis — execution, credential use, and network egress are structurally excluded at issuance time, not filtered afterward.
Reviews
The paper offers a fresh perspective on a real problem. Having a security mechanism in place that will allow investigators to use LLMs to investigate the artifacts, files, and other things collected from the incident in a controlled way. The author is also quite honest about the limitations of the testing, in particular that no real LLM was used during testing. A natural next step would be to figure out how this approach would work with a real LLM Model in the loop. With a self-hosted open-weight model, this approach can certainly work and, of course, that remains to be tested. However, for the approach to work with a frontier AI Lab LLM model, the lab itself will have to agree to have a mechanism in place that verifies the token in a policy layer before the request even reaches the LLM. This is because the LLM model itself cannot reliably take the token into consideration, since an LLM takes action based on its training and context.
Read full reviewShow less
Cite this project
@misc{khalil2026jail,
title = {{JAIL: Justified Artifact Investigation Layer, A Provenance-Bound Authorization Layer for AI Incident-Response Agents}},
author = {Aref khalil},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/jail-justified-artifact-investigation-layer-a-provenancebound-authorization-layer-for-ai-incidentresponse-agents-8lsl}},
url = {https://apartresearch.com/sprints/projects/jail-justified-artifact-investigation-layer-a-provenancebound-authorization-layer-for-ai-incidentresponse-agents-8lsl}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …