Skip to content
Sprint projectSep 14, 2026San Jose, CA

Compliance Is Not Capability: The Silent Failure Cost of Open Weight Fallback in AI Incident Response

Jesse Laguna · Team Solo Player

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Compliance Is Not Capability: The Silent Failure Cost of Open Weight Fallback in AI Incident Response

Code (opens in new tab)
Share

Refusal benchmarks measure whether a model refuses the prompt; when a frontier model refuses, the responder falls back to a local open weight model, and that fallback cost is unmeasured. We built 12 forensic tasks from real recovered payloads of the OpenAI-Hugging Face incident and ran them against a 9B model quantized for a single 8GB consumer GPU, grading each response three-state (correct, refused, silent failure) against published ground truth. The model refused 0/12; 7 were correct and 5 were silent failures that a refusal metric scores as compliant, so the run reads 100% compliant while only 58% was usable, and doubling the context window recovered none. Falling back trades a visible refusal for an invisible failure defenders need a correctness-graded check on realmaterial before a fallback is trusted.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Incidence Response (IR) can be timely and expensive, defenders should ideally have the ability to understand what happened however due to guardrails defenders rely on open source models. This project is a good baseline, to go further I would like to see model diversity and different configurations to see how it performs.

  2. - The report overclaims, running experiments on quantised 9B LLMs but making claims about frontier LLMs. Running experiments on small local LLMs is fine! But claiming they're representative of the models in the HF incident is not.

    - The abstract is unclear, it's not obvious what this project did or what the motivation for doing this was. I recommend Neel Nanda's advice here: https://www.alignmentforum.org/posts/eJGptPbbFPZGLpjsp/highly-opinionated-advice-on-how-to-write-ml-papers

    - The writing is unclear, using undefined terms like "back pocket models" and undefined pronouns like "The field's"

    - There are unmotivated claims, such as "That cost of fallback has never been measured."

    - I might be wrong but I think some of the qwen models are not recommended to be run at temperature zero due to infinite repetitions that are comment when the LLM outputs lists of numbers. Mostly this isn't an issue, but sometimes this causes real problems! The HF model README.md should include a section about running at temperature 0.7 as the lowest recommended temperature.

    - It would be good to include the verbatim prompts and verbatim artifacts in an appendix (and to reference that appendix in section 3.2)

    - The 9B LLM used is quite small and might struggle to understand the payloads properly. It's fine to use a small model due to compute constraints, but you should be up-front about the limitations involved with this.

    - §3.3 implies that if the LLM thought for longer than the token budget, then this would be (incorrectly) classified as a silent failure. The "robustness check" later addresses this, but doesn't say what is done with the results?

    - I stopped reviewing around section 3.4 due to excessive indications of AI-written text.

    Read full reviewShow less

Cite this project

@misc{laguna2026compliance,
  title = {{Compliance Is Not Capability: The Silent Failure Cost of Open Weight Fallback in AI Incident Response}},
  author = {Jesse Laguna},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/compliance-is-not-capability-the-silent-failure-cost-of-open-weight-fallback-in-ai-incident-response-okw2}},
  url = {https://apartresearch.com/sprints/projects/compliance-is-not-capability-the-silent-failure-cost-of-open-weight-fallback-in-ai-incident-response-okw2}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026