Refusals a Text-Only Scorer Cannot See: Structured Refusal Signals on a Synthetic Incident-Response Battery
Warren Smith · Team Omega
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Tests whether AI refusals can disrupt legitimate incident-response work and go undetected by evaluation systems. Across a synthetic forensic battery, one deployment refused half the tested prompts, while a text-only scorer failed to identify every refusal because the signal existed only in structured APImetadata
Reviews
This is a strong, practically useful study of a flaw in AI testing setups that is easy to miss. The key finding isn't just that one tested AI refused to help with incident response tasks. It's that those refusals were invisible to the scoring tool, which only read the AI's written reply.
In the main run, all 16 refusals came back looking like a normal, successful response, but with no text in it. The only sign of a refusal was a separate behind the scenes label saying the AI had declined. Because the scorer only looked at the text, it marked every one of them as "unclear." A rule that read that hidden label correctly identified all 16. This shows how a testing setup could easily mistake a refusal for a glitch, a missing answer, or a failure of ability.
The study also found a useful safety issue with tools. In two cases, the AI had signaled a refusal but its reply still included a request to run a tool, and the testing software ran the tool before checking whether the AI had refused. This leads to a simple, actionable recommendation: keep the behind-the-scenes stop signals, and check them before running any tools or judging a reply by its text alone.
The researchers are careful about the study's limits and open about them. The main next step is to repeat it with other AI providers, accounts, dates, and refusal detection methods, run more trials, and remove the confusing effect of the tool setup.
Overall, this is a well-done study with a lesson that can be used right away by anyone building AI tests or systems that let AI take actions.
Read full reviewShow less
Cite this project
@misc{smith2026refusals,
title = {{Refusals a Text-Only Scorer Cannot See: Structured Refusal Signals on a Synthetic Incident-Response Battery}},
author = {Warren Smith},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/refusals-a-textonly-scorer-cannot-see-structured-refusal-signals-on-a-synthetic-incidentresponse-battery-c4bt}},
url = {https://apartresearch.com/sprints/projects/refusals-a-textonly-scorer-cannot-see-structured-refusal-signals-on-a-synthetic-incidentresponse-battery-c4bt}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …