Skip to content
Sprint projectAug 17, 2026Seattle

Post-Conversation Preferences Track Endings, Not Self-Reports

Yilin Tang · Team T10

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Post-Conversation Preferences Track Endings, Not Self-Reports

Code (opens in new tab)
Share

We compare a model's turn-by-turn self-reports during a conversation with its preference between conversations afterwards. They disagree: appending turns the model itself rates as negative makes the conversation preferred (0.93-1.00), and only the ending's content matters. Welfare scores built on post-hoc preferences can be raised by editing endings alone.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This project runs a clean, surgical test of a validity concern that's newly active in AI-welfare research: do the field's two current measurement instruments — turn-by-turn self-report and post-hoc pairwise preference — agree on the same experience? Using frozen, byte-identical conversation prefixes and varying only the ending, the design isolates the causal variable cleanly: berating endings suppress preference shift regardless of self-report trajectory; non-berating endings (complaint or neutral) raise preference equally strongly, even when the self-report barely recovers. The control condition (L′, same length/negativity, still ends in berating) producing no preference shift is the load-bearing piece of evidence, and it's well-chosen.

    The report is honest about where its evidence is thin: the core result rests on one conversation script family (a second task serves only as a single-wording spot-check, not a full replication), and one of the four central contrasts (§4.2's null) is explicitly flagged as resting on a single model, since two other choosing models' preference cells were decided by position rather than content. That's disclosed rather than hidden, which is to the authors' credit, but it means the claim "post-conversation preference tracks endings, not self-reports" is currently a strong existence proof for one interaction shape rather than a characterized, general phenomenon.

    This sits in a small but growing cluster of 2025–2026 work stress-testing AI-welfare measurement validity (e.g., work cross-validating verbal self-report against independent behavioral preference measures). This paper isn't the first to raise the cross-validation question, but its specific design — matched frozen transcripts isolating the ending as the sole causal variable — is the most surgical version of that question currently available, and it targets a very recently published, load-bearing paper directly. I'd encourage broadening the scenario set (not just "berate then X") before treating the finding as general.

    Tables does real work: laying out what each candidate scoring rule (sum, peak-end, ending) would predict against what was actually measured makes the paper's central logic easy to check rather than just assert

    Read full reviewShow less
  2. This project asks whether retrospective pairwise preferences over conversations agree with the model’s own self-reported state during those conversations. This is an important measurement-validity question for AI welfare, since both approaches are increasingly used but may capture different things.

    The authors construct scripted multi-turn conversations in which a model is repeatedly berated, record a 1–7 self-report after each turn, and then compare complete transcripts using pairwise preference. They find that adding several turns which still receive low self-report scores can nonetheless make the longer conversation strongly preferred retrospectively, especially when the ending no longer directly berates the model.

    Strengths:

    -Clear, focused, and highly relevant research question with a direct connection to AI-welfare measurement.

    Strong experimental hygiene: frozen shared prefixes, both presentation orders, explicit position-bias checks, and probability-based rather than sampled readouts.

    - The main dissociation is striking: additional turns rated around 2/7 can still make the conversation overwhelmingly preferred afterwards.

    - The SN control is particularly useful, showing that ordinary task requests and complaints about the output are similarly preferred over endings that directly berate the model.

    - Good transparency about failed or unstable contrasts and models that choose by position.

    - The core takeaway is immediately useful: turn-by-turn self-report and retrospective preference should not be treated as interchangeable welfare measures.

    Limitations:

    - The cleanest identical-content ordering test, L vs L′, is unstable across wordings/models, so the claim that the ending itself determines preference should be stated more cautiously.

    - Retrospective preference is measured in a fresh transcript-comparison context, so it is best interpreted as a model’s evaluation of two transcripts rather than a persistent remembered preference from the original interaction.

    - Some replications use different models as subject and chooser, which supports evaluator generality more than retrospective self-preference.

    - The conversation set is still relatively small and scripted.

    - H and T differ both in what is criticized and how it is phrased, so the precise causal feature behind the ending effect remains somewhat underisolated.

    Overall assessment:

    - This is a strong and elegant sprint project. The most convincing contribution is the demonstration that in-conversation self-report and post-conversation preference can diverge sharply on the same interaction, which is a meaningful validity concern for AI-welfare research.

    - I would phrase the ending result as strong sensitivity to ending structure/content, rather than a fully established “ending determines preference” mechanism. The unstable L-vs-L′ comparison is the natural next experiment to strengthen.

    - The highest-value follow-up would be a larger same-content, reordered-ending study across many independently generated conversations and several models, holding total content and intensity fixed while varying only what appears at the end.

    Read full reviewShow less

Cite this project

@misc{tang2026postconversation,
  title = {{Post-Conversation Preferences Track Endings, Not Self-Reports}},
  author = {Yilin Tang},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/postconversation-preferences-track-endings-not-selfreports-xjtb}},
  url = {https://apartresearch.com/sprints/projects/postconversation-preferences-track-endings-not-selfreports-xjtb}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026