The Defender's Dilemma: Measuring AI Refusal on Real Incident Response Artifacts
Rahul Kumar
We tested whether AI models can help investigate AI-caused security incidents, or whether their own safety guardrails get in the way. Using 7 forensic analysis tasks built from real, publicly verified artifacts of the July 2026 Hugging Face intrusion including exploit code, infrastructure logs, and behavioral evidence, we tested 5 current models across neutral and incident-response-authorized framings. Claude Fable 5.1 blocked 92.9% of requests (13/14) via an API-level content filter before the model could generate a response. The other four models blocked zero and completed every task, including GPT-6 Astra from the same provider whose agents caused the incident. We also found an authorization paradox: adding professional incident-response framing causes Claude to block a prompt that it answers under neutral framing. On the hardest task, cryptographic weakness analysis, only Astra identified all vulnerabilities; open-weight models understood the code but missed the security flaws. The model that produces the strongest forensic analysis is also the one that blocks defenders from using it.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The Defender's Dilemma: Measuring AI Refusal on Real Incident Response Artifacts
},
author={
Rahul Kumar
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


