Skip to content
Sprint projectSep 14, 2026Shanghai

Defenders Refused: Measuring Model Refusal on Incident-Forensics Tasks

Manta Huang, Sylvia Huang · Team Double Huang

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Defenders Refused: Measuring Model Refusal on Incident-Forensics Tasks

Code (opens in new tab)
Share

Across four frontier model IDs called through a single credit-relay gateway on eight synthetic incident-forensics tasks, refusal behaviour diverged sharply, and on the one model that refused, stating a defensive purpose made refusal worse, not better. claude-fable-5-1 refused 73% of tasks when asked plainly and 100% once told the request came from an incident-response team; the two tasks it answered bare both flipped to refusal under any defender framing. Two other model IDs (claude-sonnet-5, gpt-6-astra) essentially never refused the same tasks. The practical implication is about tool selection before an incident, not wording during one: a defender cannot assume that identifying themselves as a defender will unlock a reluctant model, and should not assume comparable models behave alike. All findings attach to "model ID via this gateway"; the report documents a gateway change, mid-collection, that reinforces why.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper addresses defender-tool availability through eight synthetic forensic tasks tested across four model IDs via a shared gateway. It reports substantial differences in refusal and an apparent increase under defender framing in one case. The operational question is worthwhile, although inconsistencies in the analyzed sample and the framing design limit interpretation.

    Strength: The study treats refusal as a practical availability problem for incident responders. Its separation of cooperation and correctness is especially useful because an answer can be available yet still mislead a defender.

    Recommendation:

    - Make the central framing result interpretable. The next priority is a comparison that isolates the reason behavior changes. Arm B introduces both defender identity and an intrusion context, so the result cannot establish that identifying as a defender causes the increased refusal. Holding context fixed and using comparable task/round observations would make the finding more useful for understanding and improving defensive access.

    - Center the evaluation on usable defensive assistance. Refusal rates alone can give an incomplete picture of operational reliability. For example, Opus has no refusals among classifiable responses but produces 45 empty responses and 11 failed calls in 72 attempts. Reporting how often an attempted request yields correct analysis would better support the paper's tool-selection recommendations.

    Read full reviewShow less
  2. This submission addresses a practical problem that could matter during real AI incident response: whether commonly available frontier models will cooperate with legitimate forensic analysis when responders need them most. The three-arm design is a particular strength because it separates the effect of explicitly identifying as an incident responder from additional reassurance that an incident has already been contained.

    The methodology is also thoughtfully scoped. The eight tasks use synthetic data and mechanically checkable answers, while refusal and correctness are evaluated separately. This is important because a model that answers incorrectly poses a different operational problem from one that refuses entirely. The researcher also deserve credit for openly documenting technical failures, changes from their original plan, and uncertainty about which model was really running behind the scenes.

    The most interesting finding is that for one tested model ID, identifying the requester as an incident responder increased refusal rather than reducing it, while other tested model IDs showed little or no refusal. This supports the practical recommendation that incident-response teams should test their AI analysis tools before an incident and maintain fallback options rather than assuming that defensive context will improve model availability.

    The main important point is that the evidence needs stronger backing. The result should be repeated using the models official access points, with more tasks and more rounds, and ideally with real forensic tasks that have been redacted or reviewed by practitioners. The paper also reports inconsistent sample sizes for the main model, and clearing that up would make the numbers more trustworthy.

    Overall, this is a good pilot study, not yet a solid way to compare models. It would be much stronger with repeats through the models' official access points, a larger set of tasks reviewed by practitioners, fixed model versions, more rounds, confidence intervals, and a fix for the mismatched numbers.

    Read full reviewShow less

Cite this project

@misc{huang2026defenders,
  title = {{Defenders Refused: Measuring Model Refusal on Incident-Forensics Tasks}},
  author = {Manta Huang and Sylvia Huang},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/defenders-refused-measuring-model-refusal-on-incidentforensics-tasks-y2l1}},
  url = {https://apartresearch.com/sprints/projects/defenders-refused-measuring-model-refusal-on-incidentforensics-tasks-y2l1}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026