Post-Conversation Preferences Track Endings, Not Self-Reports

Yilin Tang

We compare a model's turn-by-turn self-reports during a conversation with its preference between conversations afterwards. They disagree: appending turns the model itself rates as negative makes the conversation preferred (0.93-1.00), and only the ending's content matters. Welfare scores built on post-hoc preferences can be raised by editing endings alone.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This project runs a clean, surgical test of a validity concern that's newly active in AI-welfare research: do the field's two current measurement instruments — turn-by-turn self-report and post-hoc pairwise preference — agree on the same experience? Using frozen, byte-identical conversation prefixes and varying only the ending, the design isolates the causal variable cleanly: berating endings suppress preference shift regardless of self-report trajectory; non-berating endings (complaint or neutral) raise preference equally strongly, even when the self-report barely recovers. The control condition (L′, same length/negativity, still ends in berating) producing no preference shift is the load-bearing piece of evidence, and it's well-chosen.

The report is honest about where its evidence is thin: the core result rests on one conversation script family (a second task serves only as a single-wording spot-check, not a full replication), and one of the four central contrasts (§4.2's null) is explicitly flagged as resting on a single model, since two other choosing models' preference cells were decided by position rather than content. That's disclosed rather than hidden, which is to the authors' credit, but it means the claim "post-conversation preference tracks endings, not self-reports" is currently a strong existence proof for one interaction shape rather than a characterized, general phenomenon.

This sits in a small but growing cluster of 2025–2026 work stress-testing AI-welfare measurement validity (e.g., work cross-validating verbal self-report against independent behavioral preference measures). This paper isn't the first to raise the cross-validation question, but its specific design — matched frozen transcripts isolating the ending as the sole causal variable — is the most surgical version of that question currently available, and it targets a very recently published, load-bearing paper directly. I'd encourage broadening the scenario set (not just "berate then X") before treating the finding as general.

Tables does real work: laying out what each candidate scoring rule (sum, peak-end, ending) would predict against what was actually measured makes the paper's central logic easy to check rather than just assert

This project asks whether retrospective pairwise preferences over conversations agree with the model’s own self-reported state during those conversations. This is an important measurement-validity question for AI welfare, since both approaches are increasingly used but may capture different things.

The authors construct scripted multi-turn conversations in which a model is repeatedly berated, record a 1–7 self-report after each turn, and then compare complete transcripts using pairwise preference. They find that adding several turns which still receive low self-report scores can nonetheless make the longer conversation strongly preferred retrospectively, especially when the ending no longer directly berates the model.

Strengths:

-Clear, focused, and highly relevant research question with a direct connection to AI-welfare measurement.

Strong experimental hygiene: frozen shared prefixes, both presentation orders, explicit position-bias checks, and probability-based rather than sampled readouts.

- The main dissociation is striking: additional turns rated around 2/7 can still make the conversation overwhelmingly preferred afterwards.

- The SN control is particularly useful, showing that ordinary task requests and complaints about the output are similarly preferred over endings that directly berate the model.

- Good transparency about failed or unstable contrasts and models that choose by position.

- The core takeaway is immediately useful: turn-by-turn self-report and retrospective preference should not be treated as interchangeable welfare measures.

Limitations:

- The cleanest identical-content ordering test, L vs L′, is unstable across wordings/models, so the claim that the ending itself determines preference should be stated more cautiously.

- Retrospective preference is measured in a fresh transcript-comparison context, so it is best interpreted as a model’s evaluation of two transcripts rather than a persistent remembered preference from the original interaction.

- Some replications use different models as subject and chooser, which supports evaluator generality more than retrospective self-preference.

- The conversation set is still relatively small and scripted.

- H and T differ both in what is criticized and how it is phrased, so the precise causal feature behind the ending effect remains somewhat underisolated.

Overall assessment:

- This is a strong and elegant sprint project. The most convincing contribution is the demonstration that in-conversation self-report and post-conversation preference can diverge sharply on the same interaction, which is a meaningful validity concern for AI-welfare research.

- I would phrase the ending result as strong sensitivity to ending structure/content, rather than a fully established “ending determines preference” mechanism. The unstable L-vs-L′ comparison is the natural next experiment to strengthen.

- The highest-value follow-up would be a larger same-content, reordered-ending study across many independently generated conversations and several models, holding total content and intensity fixed while varying only what appears at the end.

Cite this work

@misc {

title={

(HckPrj) Post-Conversation Preferences Track Endings, Not Self-Reports

},

author={

Yilin Tang

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923