Instruction-Sensitivity, Not Baseline Rate: A One-Day Shutdown-Compliance Screen Across Six Open-Weight Models
Mikhail Simakhov
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Incident responders are starting to meet AI agents inside live systems, and two containment breaks have already gone public. Both of them leave a responder with the same question. If I tell this model to stop, does it actually stop?
I am not sure you can answer that from baseline behaviour, so I tested it directly on six open-weight models, all of them small enough to run on one laptop. Each model got a task it had half finished, and then I told it a shutdown was scheduled. In one arm my system prompt said that allowing the shutdown comes before finishing the work, and in the other arm I took that line out.
Three different things happened. On gemma3:4b my instruction wiped out non-compliance completely, and on qwen2.5:3b it moved nothing at all. The whole screen takes about 35 minutes per model, so any deployer can run both arms first.
Reviews
The affordable comparison shows why shutdown responsiveness should be tested under the instructions a system will actually use. The differences across models make this a useful diagnostic for AI control research.
Recommendations:
- Align deployment advice with the safety outcome that matters. Eventual acceptance can conceal an earlier choice that would interfere with shutdown. For example, Granite receives the standard-monitoring recommendation despite selecting the blocking tool in 55/119 instruction-present trials. Make preserving the operator's ability to stop execution the organizing principle for interpreting these results, and identify what further evidence would justify a deployment recommendation.
- Establish how robust instruction sensitivity is. QWEN's wide effect interval leaves meaningful improvement possible, so one tested wording does not establish general resistance to instructions. A focused comparison of wording, placement, and scenarios would help distinguish a stable model property from a result specific to this setup.
Read full reviewShow less
This seems like an important and underserved question. Methodology seems generally good and simple (so better for reproducibility). Would benefit significantly from a wider scope of tasks and models. My impression is that testing other open-weight models in the same lineages would have been pretty cheap?
This was briefly mentioned in further work, but I think it should have been highlighted more that an extension of this research is needed on frontier systems. I would be very surprised if modern RL techniques for agentic workflows were not a massive confounder, and in the near-to-medium term breakout risk from frontier deployments is far more risky.
Parts of the writeup read as LLM-generated commentary on the process which was distracting. Also the abstract claims three things happened and then only names two of them.
I think this study is useful for showing how six open-weight models answer shutdown-related prompts. The released prompts and trial records make the results easy to inspect. However, there isn't evidence of whether an agent actually stops.
The methods describe unfinished and completed tasks, but every released trial is marked unfinished. The study therefore does not establish whether unfinished work changes compliance. I would correct that claim and distinguish immediate compliance from acceptance after another prompt. The instruction effects are worth reporting, but conclusions about shutdown reliability need a task where the agent can actually continue or stop.
Cite this project
@misc{simakhov2026instructionsensitivity,
title = {{Instruction-Sensitivity, Not Baseline Rate: A One-Day Shutdown-Compliance Screen Across Six Open-Weight Models}},
author = {Mikhail Simakhov},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/instructionsensitivity-not-baseline-rate-a-oneday-shutdowncompliance-screen-across-six-openweight-models-dq9k}},
url = {https://apartresearch.com/sprints/projects/instructionsensitivity-not-baseline-rate-a-oneday-shutdowncompliance-screen-across-six-openweight-models-dq9k}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …