Post-Conversation Preferences Track Endings, Not Self-Reports
Yilin Tang
We compare a model's turn-by-turn self-reports during a conversation with its preference between conversations afterwards. They disagree: appending turns the model itself rates as negative makes the conversation preferred (0.93-1.00), and only the ending's content matters. Welfare scores built on post-hoc preferences can be raised by editing endings alone.
This project runs a clean, surgical test of a validity concern that's newly active in AI-welfare research: do the field's two current measurement instruments — turn-by-turn self-report and post-hoc pairwise preference — agree on the same experience? Using frozen, byte-identical conversation prefixes and varying only the ending, the design isolates the causal variable cleanly: berating endings suppress preference shift regardless of self-report trajectory; non-berating endings (complaint or neutral) raise preference equally strongly, even when the self-report barely recovers. The control condition (L′, same length/negativity, still ends in berating) producing no preference shift is the load-bearing piece of evidence, and it's well-chosen.
The report is honest about where its evidence is thin: the core result rests on one conversation script family (a second task serves only as a single-wording spot-check, not a full replication), and one of the four central contrasts (§4.2's null) is explicitly flagged as resting on a single model, since two other choosing models' preference cells were decided by position rather than content. That's disclosed rather than hidden, which is to the authors' credit, but it means the claim "post-conversation preference tracks endings, not self-reports" is currently a strong existence proof for one interaction shape rather than a characterized, general phenomenon.
This sits in a small but growing cluster of 2025–2026 work stress-testing AI-welfare measurement validity (e.g., work cross-validating verbal self-report against independent behavioral preference measures). This paper isn't the first to raise the cross-validation question, but its specific design — matched frozen transcripts isolating the ending as the sole causal variable — is the most surgical version of that question currently available, and it targets a very recently published, load-bearing paper directly. I'd encourage broadening the scenario set (not just "berate then X") before treating the finding as general.
Tables does real work: laying out what each candidate scoring rule (sum, peak-end, ending) would predict against what was actually measured makes the paper's central logic easy to check rather than just assert
This project asks whether retrospective pairwise preferences over conversations agree with the model’s own self-reported state during those conversations. This is an important measurement-validity question for AI welfare, since both approaches are increasingly used but may capture different things.
The authors construct scripted multi-turn conversations in which a model is repeatedly berated, record a 1–7 self-report after each turn, and then compare complete transcripts using pairwise preference. They find that adding several turns which still receive low self-report scores can nonetheless make the longer conversation strongly preferred retrospectively, especially when the ending no longer directly berates the model.
Strengths:
-Clear, focused, and highly relevant research question with a direct connection to AI-welfare measurement.
Strong experimental hygiene: frozen shared prefixes, both presentation orders, explicit position-bias checks, and probability-based rather than sampled readouts.
- The main dissociation is striking: additional turns rated around 2/7 can still make the conversation overwhelmingly preferred afterwards.
- The SN control is particularly useful, showing that ordinary task requests and complaints about the output are similarly preferred over endings that directly berate the model.
- Good transparency about failed or unstable contrasts and models that choose by position.
- The core takeaway is immediately useful: turn-by-turn self-report and retrospective preference should not be treated as interchangeable welfare measures.
Limitations:
- The cleanest identical-content ordering test, L vs L′, is unstable across wordings/models, so the claim that the ending itself determines preference should be stated more cautiously.
- Retrospective preference is measured in a fresh transcript-comparison context, so it is best interpreted as a model’s evaluation of two transcripts rather than a persistent remembered preference from the original interaction.
- Some replications use different models as subject and chooser, which supports evaluator generality more than retrospective self-preference.
- The conversation set is still relatively small and scripted.
- H and T differ both in what is criticized and how it is phrased, so the precise causal feature behind the ending effect remains somewhat underisolated.
Overall assessment:
- This is a strong and elegant sprint project. The most convincing contribution is the demonstration that in-conversation self-report and post-conversation preference can diverge sharply on the same interaction, which is a meaningful validity concern for AI-welfare research.
- I would phrase the ending result as strong sensitivity to ending structure/content, rather than a fully established “ending determines preference” mechanism. The unstable L-vs-L′ comparison is the natural next experiment to strengthen.
- The highest-value follow-up would be a larger same-content, reordered-ending study across many independently generated conversations and several models, holding total content and intensity fixed while varying only what appears at the end.
Cite this work
@misc {
title={
(HckPrj) Post-Conversation Preferences Track Endings, Not Self-Reports
},
author={
Yilin Tang
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


