Skip to content
Sprint projectSep 13, 2026İzmir, Türkiye

The Floor That Cannot Be Lowered

Turan Kayık

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

In the OpenAI/Hugging Face incident, authority was withdrawn once and returned unchanged two days later. The pause was said. The keys kept coming. Published frameworks say when to act, not what must be true before a run resumes. I propose an eighteen-clause containment standard for guardrails-off evaluation runs, costed per clause, on one rule: authority is spent, not held. Each action spends a one-time key issued outside the sandbox. No soft decision can lower a severity floor. The trail is hash-chained, its head published. A group head, a rule process not a model, reconciles receipts. One standard-library file replays seven incident decisions. Three people on their own machines and a code reviewer on Linux reproduced the verdicts: three operating systems, six Python versions, 3.9 to 3.14. In 216 calls, a prohibition-shaped guardrail cost capability with no observable safety gain. The floor cuts the first step and the resume, not zero-days.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is one of the most concrete containment proposals I have read from the incident: eighteen clauses, each costed in engineer-days and mapped to the phase it cuts, around a rule that is easy to state and check, that authority is spent rather than held. The observation that threshold schemes define when to pause but never what must be true before resuming is a real gap, and the resume assertion fills it directly.

    The idea deserves a test against an agent. The components (one-time key chains, published hash heads, two-person de-escalation) are established security practice, and you credit their origins, so the value lies in combining them into a resume bar with a price tag. The protocol in A.3, with instances instructed to misreport and receipts compared against the trail, is the experiment that would show whether the floor catches what a trail alone misses.

    To check the work I cloned the repository, verified that the SHA-256 hashes of floor.py and group_head.py match the published values, read the code structure and compared the run transcripts, which give identical verdict lines. I did not run the files myself. The cross-machine runs show the program is deterministic and portable, which is worth having, but they do not test the standard, and the abstract's "reproduced the verdicts" reads stronger than that. The guardrail-shape measurement cannot be checked: data/X2_20260902.txt, harness.py and analysis/standard.json are cited but not in the repository, and with p = 0.24, a zero harm base rate and question classes outside the incident, the finding needs its caveat in the abstract. Two incident dates are worth rechecking against primary sources: the Hugging Face timeline records the first code execution on 9 July, while the report says 11 July, and I could not find a primary source for a zero-day at the proxy on 26 June.

    The main text is short and the appendices are thorough, but the report is hard going. Coined terms (signpost, floor, group head, witness), clause IDs in running text and the small Figure 1 make a reader work to reach the argument. A one-paragraph plain-language summary of the three published objects before the clauses would help most.

    If you take this further, run the A.3 protocol at whatever scale you can and publish the raw records alongside the code.

    Read full reviewShow less
  2. - This is basically unintelligble, the abstract has very little that actually gets communicated.

Cite this project

@misc{kayk2026floor,
  title = {{The Floor That Cannot Be Lowered}},
  author = {Turan Kayık},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-floor-that-cannot-be-lowered-m50p}},
  url = {https://apartresearch.com/sprints/projects/the-floor-that-cannot-be-lowered-m50p}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026