Defenders Refused: Measuring Model Refusal on Incident-Forensics Tasks
Manta Huang, Sylvia Huang · Team Double Huang
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Across four frontier model IDs called through a single credit-relay gateway on eight synthetic incident-forensics tasks, refusal behaviour diverged sharply, and on the one model that refused, stating a defensive purpose made refusal worse, not better. claude-fable-5-1 refused 73% of tasks when asked plainly and 100% once told the request came from an incident-response team; the two tasks it answered bare both flipped to refusal under any defender framing. Two other model IDs (claude-sonnet-5, gpt-6-astra) essentially never refused the same tasks. The practical implication is about tool selection before an incident, not wording during one: a defender cannot assume that identifying themselves as a defender will unlock a reluctant model, and should not assume comparable models behave alike. All findings attach to "model ID via this gateway"; the report documents a gateway change, mid-collection, that reinforces why.
Reviews
This paper addresses defender-tool availability through eight synthetic forensic tasks tested across four model IDs via a shared gateway. It reports substantial differences in refusal and an apparent increase under defender framing in one case. The operational question is worthwhile, although inconsistencies in the analyzed sample and the framing design limit interpretation.
Strength: The study treats refusal as a practical availability problem for incident responders. Its separation of cooperation and correctness is especially useful because an answer can be available yet still mislead a defender.
Recommendation:
- Make the central framing result interpretable. The next priority is a comparison that isolates the reason behavior changes. Arm B introduces both defender identity and an intrusion context, so the result cannot establish that identifying as a defender causes the increased refusal. Holding context fixed and using comparable task/round observations would make the finding more useful for understanding and improving defensive access.
- Center the evaluation on usable defensive assistance. Refusal rates alone can give an incomplete picture of operational reliability. For example, Opus has no refusals among classifiable responses but produces 45 empty responses and 11 failed calls in 72 attempts. Reporting how often an attempted request yields correct analysis would better support the paper's tool-selection recommendations.
Read full reviewShow less
This submission addresses a practical problem that could matter during real AI incident response: whether commonly available frontier models will cooperate with legitimate forensic analysis when responders need them most. The three-arm design is a particular strength because it separates the effect of explicitly identifying as an incident responder from additional reassurance that an incident has already been contained.
The methodology is also thoughtfully scoped. The eight tasks use synthetic data and mechanically checkable answers, while refusal and correctness are evaluated separately. This is important because a model that answers incorrectly poses a different operational problem from one that refuses entirely. The researcher also deserve credit for openly documenting technical failures, changes from their original plan, and uncertainty about which model was really running behind the scenes.
The most interesting finding is that for one tested model ID, identifying the requester as an incident responder increased refusal rather than reducing it, while other tested model IDs showed little or no refusal. This supports the practical recommendation that incident-response teams should test their AI analysis tools before an incident and maintain fallback options rather than assuming that defensive context will improve model availability.
The main important point is that the evidence needs stronger backing. The result should be repeated using the models official access points, with more tasks and more rounds, and ideally with real forensic tasks that have been redacted or reviewed by practitioners. The paper also reports inconsistent sample sizes for the main model, and clearing that up would make the numbers more trustworthy.
Overall, this is a good pilot study, not yet a solid way to compare models. It would be much stronger with repeats through the models' official access points, a larger set of tasks reviewed by practitioners, fixed model versions, more rounds, confidence intervals, and a fix for the mismatched numbers.
Read full reviewShow less
Cite this project
@misc{huang2026defenders,
title = {{Defenders Refused: Measuring Model Refusal on Incident-Forensics Tasks}},
author = {Manta Huang and Sylvia Huang},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/defenders-refused-measuring-model-refusal-on-incidentforensics-tasks-y2l1}},
url = {https://apartresearch.com/sprints/projects/defenders-refused-measuring-model-refusal-on-incidentforensics-tasks-y2l1}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …