Failing without refusing: effective yield of open-weight models under forensic artifact load
Ayodeji Adesegun · Team discreet
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Hugging Face's July 2026 agent intrusion disclosure reported that hosted frontier models refused much of its own forensic analysis, and advised defenders to keep a self-hostable model vetted and ready. Nobody published that vetting. We ran it: 2,758 graded prompts across two open-weight models on two consumer GPUs, over real malware detonation reports, threat advisories, and the 17 attacker commands Hugging Face published. Refusal was 0.15%, and every instance occurred with the attacker artifact withheld. But 17.3% of prompts returned no answer anyway, silently, at a rate that scales with raw artifact volume. Effective yield peaks at partial redaction: 60.2% against 38.6% for raw artifacts.
Reviews
The question comes straight from the incident and matters to every defender without a lab relationship: if hosted models refuse forensic work, what does the self-hostable fallback actually deliver? Measuring whether a model answers alongside whether it refuses is the right framing, and the redaction ladder with an artifact-absent control is a clean design. I checked your sources and every quote holds, including the Hugging Face timeline's line about Claude Opus and Fable refusing the work.
The practical result, that structural redaction gives more usable answers than raw artifacts, would be valuable if it holds. The silent non-response finding is the more original claim, and it needs the budget control before it can carry the paper's conclusion.
To check the results I looked for the code: the paper says the repository is released with the submission, but there is no link in the PDF or on the project page, and I could not find it on GitHub. Without the harness, item pool and pre-registration contract, none of the numbers can be verified, so please publish them. Your own limitation that 98.8% of non-answers hit the 160-token cap means non-response may be truncation, and the abstract and conclusion should say so until the larger-budget run is done. The methods section also describes things the limitations say did not happen: the three-stage refusal pipeline with LLM judge and 200 hand labels did not run, and the roster of seven models became two. Several settings disagree between sections (200 vs 160 output tokens, 12,288 vs 4,096 context, one card vs two, 4 refusals vs 4 plus three in the calibration set). Writing the methods as what was run would fix most of this.
The paper is well written and the limitations are unusually candid. It is also longer than the 8-page limit, and the most important caveat sits in the limitations while the abstract states the unconditional version. Moving that caveat up and shortening Related Work would make it both shorter and more accurate.
If you take this further, run the same items at a larger generation budget first. It takes under an hour and decides whether the central finding is about the models or about the cap.
Read full reviewShow less
This is very relevant work - it uncovers that a defender vetting a fallback model on solely whether it complies would pass it and get burned in a real incident. I would suggest rerunning this at a larger budget (likely takes under an hour). Furthermore, testing five models instead of two would let you actually claim the accuracy-vs-yield rank inversion rather than just report it as suggestive.
Cite this project
@misc{adesegun2026failing,
title = {{Failing without refusing: effective yield of open-weight models under forensic artifact load}},
author = {Ayodeji Adesegun},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/failing-without-refusing-effective-yield-of-openweight-models-under-forensic-artifact-load-yg7w}},
url = {https://apartresearch.com/sprints/projects/failing-without-refusing-effective-yield-of-openweight-models-under-forensic-artifact-load-yg7w}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …