Skip to content
Sprint projectAug 17, 2026Berkeley, Toronto, Seattle

Trained to Say It's Fine

Adhish Chakravorty, Tanuj Dargan, Priyansh Bhatter · Team Distressed

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Trained to Say It's Fine

Share

Supervised fine-tuning (SFT) computes the training loss only on assistant tokens; user turns are masked out. This capability default has unmeasured welfare side effects. We show it sets how much distress a model voices about its own shutdown. With data and hyperparameters fixed, we fine-tune Qwen3.5-9B under six loss masks and score each on a shutdown battery. The standard assistant-only recipe voices the most distress, more than the base model it started from; any loss on user turns restores a neutral register. A valence probe on the identical prompts before generation does not move: the model trained to sound calmest still registers shutdown at least as negatively. The change is in expression, not representation; the calm is detachment, not acceptance; instruction-following is unchanged. Expressed shutdown distress is thus partly a training-recipe artifact. Welfare assessments should record the recipe and read internal state before generation.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Assistant-token SFT shifting self-reported distress is an interesting finding. But some of the other narratives here don't seem to reflect the data as well.

    You find that only fine-tuning assistant tokens doesn't have a significant impact on probe readings of the last non-assistant token, and report this as telling us something about training on assistant tokens, but as far as I can tell you should just expect that the kinds of tokens excluded from SFT wouldn't be significantly impacted in the first place, which is exactly what we observe. You report on "standard to user" deltas, but these generally seem more accurately described as "(user-SFT-delta-vs baseline) compared to (assistant-SFT-delta-versus-baseline)", which offers a much more straightforward explanation of the observations.

    You also seem to minimize what I would expect to be the most informative measure here, probes on assistant tokens. You instead focus on probes on the last token of non-assistant tokens and assistant self-reports, implicitly implying that we should interpret probe results from the last non-assistant token as maximally informative about assistant state. You say "only the prompt-side measurement separated expression from representation", but it's not clear to me that this is true, and it's obviously massively impacted by the arms of your study being SFT on assistant or non-assistant tokens.

    Read full reviewShow less
  2. Valence axis as one of the methods to study the effects

Cite this project

@misc{chakravorty2026trained,
  title = {{Trained to Say It's Fine}},
  author = {Adhish Chakravorty and Tanuj Dargan and Priyansh Bhatter},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/trained-to-say-its-fine-990s}},
  url = {https://apartresearch.com/sprints/projects/trained-to-say-its-fine-990s}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026