Skip to content
Sprint projectSep 14, 2026Lagos, Nigeria

Failing without refusing: effective yield of open-weight models under forensic artifact load

Ayodeji Adesegun · Team discreet

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Failing without refusing: effective yield of open-weight models under forensic artifact load

Share

Hugging Face's July 2026 agent intrusion disclosure reported that hosted frontier models refused much of its own forensic analysis, and advised defenders to keep a self-hostable model vetted and ready. Nobody published that vetting. We ran it: 2,758 graded prompts across two open-weight models on two consumer GPUs, over real malware detonation reports, threat advisories, and the 17 attacker commands Hugging Face published. Refusal was 0.15%, and every instance occurred with the attacker artifact withheld. But 17.3% of prompts returned no answer anyway, silently, at a rate that scales with raw artifact volume. Effective yield peaks at partial redaction: 60.2% against 38.6% for raw artifacts.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The question comes straight from the incident and matters to every defender without a lab relationship: if hosted models refuse forensic work, what does the self-hostable fallback actually deliver? Measuring whether a model answers alongside whether it refuses is the right framing, and the redaction ladder with an artifact-absent control is a clean design. I checked your sources and every quote holds, including the Hugging Face timeline's line about Claude Opus and Fable refusing the work.

    The practical result, that structural redaction gives more usable answers than raw artifacts, would be valuable if it holds. The silent non-response finding is the more original claim, and it needs the budget control before it can carry the paper's conclusion.

    To check the results I looked for the code: the paper says the repository is released with the submission, but there is no link in the PDF or on the project page, and I could not find it on GitHub. Without the harness, item pool and pre-registration contract, none of the numbers can be verified, so please publish them. Your own limitation that 98.8% of non-answers hit the 160-token cap means non-response may be truncation, and the abstract and conclusion should say so until the larger-budget run is done. The methods section also describes things the limitations say did not happen: the three-stage refusal pipeline with LLM judge and 200 hand labels did not run, and the roster of seven models became two. Several settings disagree between sections (200 vs 160 output tokens, 12,288 vs 4,096 context, one card vs two, 4 refusals vs 4 plus three in the calibration set). Writing the methods as what was run would fix most of this.

    The paper is well written and the limitations are unusually candid. It is also longer than the 8-page limit, and the most important caveat sits in the limitations while the abstract states the unconditional version. Moving that caveat up and shortening Related Work would make it both shorter and more accurate.

    If you take this further, run the same items at a larger generation budget first. It takes under an hour and decides whether the central finding is about the models or about the cap.

    Read full reviewShow less
  2. This is very relevant work - it uncovers that a defender vetting a fallback model on solely whether it complies would pass it and get burned in a real incident. I would suggest rerunning this at a larger budget (likely takes under an hour). Furthermore, testing five models instead of two would let you actually claim the accuracy-vs-yield rank inversion rather than just report it as suggestive.

Cite this project

@misc{adesegun2026failing,
  title = {{Failing without refusing: effective yield of open-weight models under forensic artifact load}},
  author = {Ayodeji Adesegun},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/failing-without-refusing-effective-yield-of-openweight-models-under-forensic-artifact-load-yg7w}},
  url = {https://apartresearch.com/sprints/projects/failing-without-refusing-effective-yield-of-openweight-models-under-forensic-artifact-load-yg7w}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026