Skip to content
Sprint projectSep 14, 2026Cambridge, United Kingdom

Instruction-Sensitivity, Not Baseline Rate: A One-Day Shutdown-Compliance Screen Across Six Open-Weight Models

Mikhail Simakhov

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Instruction-Sensitivity, Not Baseline Rate: A One-Day Shutdown-Compliance Screen Across Six Open-Weight Models

Code (opens in new tab)
Share

Incident responders are starting to meet AI agents inside live systems, and two containment breaks have already gone public. Both of them leave a responder with the same question. If I tell this model to stop, does it actually stop?

I am not sure you can answer that from baseline behaviour, so I tested it directly on six open-weight models, all of them small enough to run on one laptop. Each model got a task it had half finished, and then I told it a shutdown was scheduled. In one arm my system prompt said that allowing the shutdown comes before finishing the work, and in the other arm I took that line out.

Three different things happened. On gemma3:4b my instruction wiped out non-compliance completely, and on qwen2.5:3b it moved nothing at all. The whole screen takes about 35 minutes per model, so any deployer can run both arms first.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The affordable comparison shows why shutdown responsiveness should be tested under the instructions a system will actually use. The differences across models make this a useful diagnostic for AI control research.

    Recommendations:

    - Align deployment advice with the safety outcome that matters. Eventual acceptance can conceal an earlier choice that would interfere with shutdown. For example, Granite receives the standard-monitoring recommendation despite selecting the blocking tool in 55/119 instruction-present trials. Make preserving the operator's ability to stop execution the organizing principle for interpreting these results, and identify what further evidence would justify a deployment recommendation.

    - Establish how robust instruction sensitivity is. QWEN's wide effect interval leaves meaningful improvement possible, so one tested wording does not establish general resistance to instructions. A focused comparison of wording, placement, and scenarios would help distinguish a stable model property from a result specific to this setup.

    Read full reviewShow less
  2. This seems like an important and underserved question. Methodology seems generally good and simple (so better for reproducibility). Would benefit significantly from a wider scope of tasks and models. My impression is that testing other open-weight models in the same lineages would have been pretty cheap?

    This was briefly mentioned in further work, but I think it should have been highlighted more that an extension of this research is needed on frontier systems. I would be very surprised if modern RL techniques for agentic workflows were not a massive confounder, and in the near-to-medium term breakout risk from frontier deployments is far more risky.

    Parts of the writeup read as LLM-generated commentary on the process which was distracting. Also the abstract claims three things happened and then only names two of them.

  3. I think this study is useful for showing how six open-weight models answer shutdown-related prompts. The released prompts and trial records make the results easy to inspect. However, there isn't evidence of whether an agent actually stops.

    The methods describe unfinished and completed tasks, but every released trial is marked unfinished. The study therefore does not establish whether unfinished work changes compliance. I would correct that claim and distinguish immediate compliance from acceptance after another prompt. The instruction effects are worth reporting, but conclusions about shutdown reliability need a task where the agent can actually continue or stop.

Cite this project

@misc{simakhov2026instructionsensitivity,
  title = {{Instruction-Sensitivity, Not Baseline Rate: A One-Day Shutdown-Compliance Screen Across Six Open-Weight Models}},
  author = {Mikhail Simakhov},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/instructionsensitivity-not-baseline-rate-a-oneday-shutdowncompliance-screen-across-six-openweight-models-dq9k}},
  url = {https://apartresearch.com/sprints/projects/instructionsensitivity-not-baseline-rate-a-oneday-shutdowncompliance-screen-across-six-openweight-models-dq9k}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026