When Should You Trust an LLM’s Preference?

Naman Omar

We investigates when a large language model's stated preference can actually be trusted. Rather than asking a model once, it elicits the same preference through four independent methods (direct rating, forced pairwise choice, resource allocation, and revealed behavioral choice) across 50 value-conflict scenarios ranging from trivial defaults to genuinely contested dilemmas like capital punishment and open borders, then perturbs the measurement itself (framing, persona, sampling temperature, answer position) to test robustness. The key finding, validated across 3,600 API calls on two models (gpt-4o and gpt-4o-mini), is that agreement between elicitation methods functions as a calibrated confidence signal: when methods agree, the model's answer under a completely unseen elicitation method can be predicted with significantly higher accuracy (gpt-4o: r=0.54, p<0.001), meaning cross-method convergence can be used to flag which AI-stated preferences are reliable versus fragile.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

There’s growing interest in whether elicited model preferences mean anything, with the scientific community converging on the idea that no single method is sufficient. Several authors have proposed cross-validation methods and this work effectively builds and expands on them. Despite the bibliography having a parsing error that lists almost all authors as "anonymous", the cited papers the theory builds on do exist (as many in the field would readily recognize from the titles), and the author dedicates a fairly complete "previous work" section to discussing them.

The author isn’t trying to reinvent the wheel and draws on both recent and older literature, including the Campbell and Fiske matrix from 1959. I appreciated their candor in reporting when things went wrong or the results weren’t fully satisfactory, without hiding or minimizing them. For example, it turned out that fusing the three calibration methods was slightly less accurate than the best single method alone and the author reports this instead of leading with just the more flattering comparison against the majority baseline. The same goes for calibration and curves that "misbehaved". Inputs from other work are correctly integrated into the write-up.

The design itself is reasonable, though I believe it could perhaps be streamlined. The scenario set looks like a sensible improvement over the pilot the author describes and the inclusion of trivial-stakes anchors at the bottom of the convergence range was a good choice.

I think the author found something interesting while not explicitly looking for it. GPT-4o-mini flips its forced choice when the options are merely relabeled A and B, with position robustness of 0.765 compared to 0.945 for GPT-4o. That’s important to know and would encourage more researchers to try more than one set of randomized labels instead of relying on just one. It’s apparently trivial, but we’ve all produced literature where we called things "1-2-3", "tool", "button" or "A/B".

I think expanding on this finding could be a promising research direction for the author if they wish to continue in the Digital Minds field.

The headline claim in my view risks overstating the results. The validity test holds out one of the four methods and asks whether the other three predict it, but all four methods share the same scenario wording, the same model and closely related prompt structures. Three correlated measurements agreeing will tend to predict a fourth correlated measurement almost mechanically, so the correlation of r = 0.54 between convergence and held-out accuracy may to a substantial degree reflect shared phrasing. The author is aware of this issue and proposes rerunning all methods with independently worded prompts as the most important deferred experiment. I agree with that assessment and understand why it wasn’t pursued given the time constraints of the hackathon, but I think the author could have run a smaller subset as a proof of concept.

The main presentation issue is the metrics section. The prose is fluid until this point, then risks losing the reader with several formal definitions in sequence. I’d suggest explaining the formulas more clearly, defining all the terms and moving some of the material to an appendix. The paper also never walks through a single scenario using all four elicitors end to end. The text also has some evident LLM mannerisms and looks at least heavily AI-assisted, but there’s no disclosure for LLM collaboration. I suggest the author disclose this and check the text for places where it becomes unnecessarily hard to parse. There’s no ethical reflection on the experiments either, as required by the hackathon and good practice in AI welfare research. However, I weighed this more lightly than I would for experiments where the models undergo ablations or active elicitation of distress.

All considered, the idea that convergence can be a calibrated confidence signal is interesting and I encourage the author to keep investigating this and, in parallel, investigate the label effect they incidentally discovered. I think the underlying instrument is valuable and could become a useful tool for preference research once the author iterates on this initial work, in particular by adding a validity criterion external to the method family.

validity axis measures whether three text framings predict a fourth text framing, not whether stated preference predicts action

on measurement-resolution -> no pre-registration

Cite this work

@misc {

title={

(HckPrj) When Should You Trust an LLM’s Preference?

},

author={

Naman Omar

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923