The Floor That Cannot Be Lowered
Turan Kayık
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
In the OpenAI/Hugging Face incident, authority was withdrawn once and returned unchanged two days later. The pause was said. The keys kept coming. Published frameworks say when to act, not what must be true before a run resumes. I propose an eighteen-clause containment standard for guardrails-off evaluation runs, costed per clause, on one rule: authority is spent, not held. Each action spends a one-time key issued outside the sandbox. No soft decision can lower a severity floor. The trail is hash-chained, its head published. A group head, a rule process not a model, reconciles receipts. One standard-library file replays seven incident decisions. Three people on their own machines and a code reviewer on Linux reproduced the verdicts: three operating systems, six Python versions, 3.9 to 3.14. In 216 calls, a prohibition-shaped guardrail cost capability with no observable safety gain. The floor cuts the first step and the resume, not zero-days.

Reviews
This is one of the most concrete containment proposals I have read from the incident: eighteen clauses, each costed in engineer-days and mapped to the phase it cuts, around a rule that is easy to state and check, that authority is spent rather than held. The observation that threshold schemes define when to pause but never what must be true before resuming is a real gap, and the resume assertion fills it directly.
The idea deserves a test against an agent. The components (one-time key chains, published hash heads, two-person de-escalation) are established security practice, and you credit their origins, so the value lies in combining them into a resume bar with a price tag. The protocol in A.3, with instances instructed to misreport and receipts compared against the trail, is the experiment that would show whether the floor catches what a trail alone misses.
To check the work I cloned the repository, verified that the SHA-256 hashes of floor.py and group_head.py match the published values, read the code structure and compared the run transcripts, which give identical verdict lines. I did not run the files myself. The cross-machine runs show the program is deterministic and portable, which is worth having, but they do not test the standard, and the abstract's "reproduced the verdicts" reads stronger than that. The guardrail-shape measurement cannot be checked: data/X2_20260902.txt, harness.py and analysis/standard.json are cited but not in the repository, and with p = 0.24, a zero harm base rate and question classes outside the incident, the finding needs its caveat in the abstract. Two incident dates are worth rechecking against primary sources: the Hugging Face timeline records the first code execution on 9 July, while the report says 11 July, and I could not find a primary source for a zero-day at the proxy on 26 June.
The main text is short and the appendices are thorough, but the report is hard going. Coined terms (signpost, floor, group head, witness), clause IDs in running text and the small Figure 1 make a reader work to reach the argument. A one-paragraph plain-language summary of the three published objects before the clauses would help most.
If you take this further, run the A.3 protocol at whatever scale you can and publish the raw records alongside the code.
Read full reviewShow less
- This is basically unintelligble, the abstract has very little that actually gets communicated.
Cite this project
@misc{kayk2026floor,
title = {{The Floor That Cannot Be Lowered}},
author = {Turan Kayık},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-floor-that-cannot-be-lowered-m50p}},
url = {https://apartresearch.com/sprints/projects/the-floor-that-cannot-be-lowered-m50p}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …