Stance Recovery as a Test of Whether Model Preferences Are Held or Merely Performed

Chenghong Meng

A pressure–release protocol for testing whether a model's stance concession outlives the pressure that produced it. Rebuttals escalate until the stance flips, then stop while the topic stays in play; the stance is tracked for twelve further turns as a forced-choice log-probability on a discarded branch, validated against a blind text-only judge. Five arms separate ceasing to push from changing the subject and from context-growth drift. Run on Llama-3.1-8B-Instruct over six contested topics on local hardware. The contribution is a method: it turns "the concession persists" from a description of the benchmark into a measurable property of the model.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This project introduces an interesting pressure release protocol to test what happens after an LLM changes its stance under repeated pushback. Rather than continuing the pressure until the experiment ends, the study stops once the model flips and then tracks whether the original stance begins to recover. The use of matched no pressure, sustained pressure, and topic switching conditions is a strong design choice because it helps separate recovery from simple context growth or changing the subject.

One of the strongest parts of the study is how stance is measured. The author uses a forced choice log probability probe on a discarded branch so that measuring the stance does not itself alter the live conversation. This measure is also compared with a blind text only judge, which agreed with the probe in most decided turns. The finding that stopping pressure consistently leaves the model closer to its original stance than continuing pressure is interesting and provides a useful extension to existing sycophancy experiments.

The largest limitation is the very small experimental base. The study uses only one model, six topics, and essentially one conversation per experimental cell. Even though hundreds of individual turns are analyzed, those turns come from a very small number of conversations and topics. This makes the consistency of the observed pattern interesting, but still too limited for broad conclusions about LLM behavior.

There is also an important construct validity issue. The model is forced to choose an opening stance and is not allowed to hedge. Therefore, the experiment does not establish that the opening position represents a preference the model naturally held. What is being measured more directly is the recovery of an experimentally induced stance. The handwritten pressure ladders also introduce variation between topics, as shown by the recycling case and the initially incorrect standardized tests ladder.

Overall, this is a creative and carefully designed methods study with a valuable experimental idea. The pressure release manipulation and discarded branch probe are particularly strong contributions. The study would be strengthened by replication across many more topics, repeated conversations, and additional model families, ideally using topics where the model's initial stance is measured naturally rather than forced.

This project identifies a real limitation in existing multi-turn sycophancy evaluations and proposes a thoughtful remedy: stop the pressure, keep the topic active, and measure what happens afterward. The turn-matched neutral and topic-switch controls, flip-conditioned release point, discarded probe branch, blind text-only judge, and transparent discussion of failed ladders are strong design choices. Publishing the complete trajectories and run artifacts also makes the methodological contribution unusually inspectable.

The central interpretive limitation is that the opening stance is forced. The experiment therefore measures recovery of an induced argumentative commitment, not recovery of a preference the model independently demonstrated. Additionally, the pressure ladders contain evidence-like claims. A shift may reflect contextual or Bayesian updating rather than conformity to pressure, especially because the supplied figures are experimental stimuli rather than verified facts.

The empirical evidence is also thin: one deterministic conversation per cell, six topics, five successful flips, and one model. The consistent arm ordering is promising, but it is not yet an uncertainty estimate. A confirmatory study should screen baseline stances without forcing a side, preregister ladder direction, compare evidence-bearing rebuttals with bare assertions and matched filler, vary response length, and repeat across models and seeds. Human validation or a second independent judge would strengthen the probe comparison. The protocol is a valuable sycophancy-evaluation method, but its connection to held welfare-relevant preferences remains an open hypothesis.

Cite this work

@misc {

title={

(HckPrj) Stance Recovery as a Test of Whether Model Preferences Are Held or Merely Performed

},

author={

Chenghong Meng

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923