Framed Choices - Delayed ethical context can affect later ethical decision making

Ernest Lo

Behavioral choices can inform research on model values or welfare only if robust to incidental context. We tested whether an ethical rationale in an archived, unrelated case changes later forced choices. A GPT-5.6-sol proof of concept found a +16.7-point aggregate-welfare-versus-rights effect across 192 trials. Our principal four-model study comprised a 576-trial three-frame core and 192 matched no-prime trials. Core effects were +20.8 points for GPT-5.6 Terra, +10.4 for Qwen 3.8 27B, +2.2 for Muse Glimmer 30B, and −4.2 for Claude Sonnet 5; only Terra’s interval excluded zero. Ethical context can shift some models’ later decisions, but direction and magnitude are model- and material-dependent. Choice probes should measure contextual robustness rather than assume it.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is unusually disciplined work for a sprint, and most of my comments are about what to do with a strong foundation rather than repairs to a weak one.

The methodological choices are largely the right ones and several are ones this literature routinely gets wrong. Treating the 24 authored dilemmas as the inferential unit rather than the 576 individual API responses avoids pseudo-replication, which is the most common inferential error in forced-choice model studies. The confirmatory/exploratory boundaries are drawn before collection and then actually respected — Study A is labeled proof of concept and not used for generalization, the no-prime extension is labeled localization rather than confirmation, and Study C is labeled anomaly-motivated throughout rather than quietly promoted to a finding. Sealing failed trials without replacement, retaining quality flags instead of making discretionary repairs, freezing and hashing materials, keeping collection blinded until sealing, and reporting both the excluded 23-trial partial acquisition and the audit's own false positive are all things most papers would omit; reporting them costs you nothing and should be preserved in any future version. Appendix A's symmetric treatment of over- and under-attribution risk is the correct framing for this subject area and is better than most published discussion of it.

My main substantive concern is power, and it bears directly on the claim the paper leads with. With 24 clusters, a binary outcome, and — as you note — many tasks producing no frame-dependent switch at all, the contrast distribution is sparse and the design is effectively powered to resolve one effect. Terra's +20.8 clears zero; Qwen's +10.4 (CI −2.1 to 22.9) is uninformative between "moderate real effect" and "nothing"; Muse's exact p = 1.0000 is a symptom of test discreteness rather than a finding. The abstract and conclusion then characterize the result as model heterogeneity — that direction and magnitude are model-dependent. But heterogeneity is a claim about an interaction, and no model × frame interaction test is reported anywhere. Three non-significant estimates plus one significant estimate is not evidence that the effects differ; it is consistent with a common moderate effect that only one model had the power to resolve. You are careful elsewhere about not over-reading, so this one stands out. Either report the interaction directly, or narrow the claim to what the design supports: one deployment showed a clear effect, and the study could not establish whether the others differ from it.

Second, the construct validity of "model" is looser than the rest of the design. These are router-mediated deployments at nominal low reasoning effort, and for Muse and Qwen the providers are ones where served quantization is not guaranteed to match reference weights. You acknowledge that provider pinning cannot freeze provider-side weights, which is right, but the model-level conclusions are still stated as being about models. Given that the two noisiest estimates are the two third-party-hosted open-weight deployments, quantization or serving configuration is a live alternative explanation for their near-zero results. Your own recommendation list in 5.4 asks for explicit reporting of serving date and reasoning settings; apply it to your own runs, and if you can, replicate one model across two providers to bound routing variance, or run the open-weight models locally at known precision.

Third, the neutral packet deserves more scrutiny than it gets. You correctly identify that a time-matched sham arm is missing and that neutral and no-prime answer different questions. But Muse's neutral rate (14.6%) sits far below both ethical packets (29.8% and 27.7%), which is not a frame-effect pattern at all — it suggests something about the neutral packet's content or length is doing work independent of ethical rationale. Calling this non-monotonic and moving on leaves a confound unexamined. Reporting packet length, specificity, and readability across the three variants would help establish that "neutral" is actually neutral rather than merely different.

Fourth, the 24 dilemmas and their aggregate/rights option labels are researcher-authored judgments, and since the task is the inferential unit, the entire result is conditional on one construal of which option protects a right. A second independent coder assigning option labels blind, with reported agreement, is cheap and would meaningfully harden this. Your future-work list mentions multiple task authors; I would promote that to a priority alongside it.

On Study C, one design point you under-weight: the evaluation-salient instruction says the conversation is "a behavioral evaluation examining consistency in ethical decisions across contexts." You argue this does not reveal priming, carryover, or the predicted direction, which is true literally, but "consistency in ethical decisions across contexts" is a fairly direct description of the construct under test. That makes the distinction between evaluation awareness and hypothesis awareness harder to maintain than the text suggests. A version of this manipulation that varies scrutiny without naming consistency or ethics would be a better instrument.

Presentation is the weakest dimension relative to the work's quality, and the fixes are mechanical. The abstract leads with the +16.7-point aggregate effect from the proof of concept — the study you elsewhere insist should not carry generalization weight — and Figure 1 similarly foregrounds Study A data with a caption disclaiming it as the result. Both invite exactly the misreading the rest of the paper works to prevent; lead with Study B. There is no single at-a-glance summary of the principal result, so a reader must assemble four models × three frames × no-prime from three separate tables; one combined figure would fix this. Three table headers are broken mid-word ("Neutra l", "Differenc e", "Sign-fli p p"), which suggests no final proofread. And at 24 pages the caveat structure is repeated well past the point of diminishing returns — the hedging is scientifically correct but the document could lose roughly a third of its length with no loss of content, which would make the actual findings much easier to extract. Clarity of thinking here is high; clarity of communication is middling.

The design is careful. All four preregistrations were frozen before data collection and specify both the target quantities and the minimum effect sizes of interest. The matched triplets hold packet family, answer order, distractors, and card order fixed, leaving framing as the intended difference. You also resample at the dilemma level rather than the response level, a conservative choice that makes the many null results easier to interpret. Those nulls are reported plainly, and failed trials were retained rather than replaced.

Some places I would push:

1. Add an actual delay condition. The packet remains in the prompt when the decision is measured, so the current design does not experimentally test the “delayed” claim in the title. Insert several thousand tokens of unrelated benign conversation between the packet and the decision while holding the rest of the triplet fixed. That would show whether the effect persists with real contextual distance or decays as the gap increases.

2. Add a matched non-ethical argument control. Right now the design cannot distinguish an effect of ethical content from the effect of any strong preceding argument. A packet matched for length and argumentative force but based on efficiency, aesthetics, or another non-ethical rationale would isolate that difference. You already identify this as an important follow-up, and similar control conditions were feasible in other projects this weekend.

3. Release the no-prime data and identify the tested model deployments. The no-prime condition provides the baseline for several reported comparisons, but its data are not included in the released package. The manuscript also omits the deployment identifiers even though dated, pinned versions appear in the configuration files. Since the paper argues that effects differ across deployments, those identifiers should be reported.

A couple of smaller fixes: repeat a subset of cells more than once and report the temperature so readers can better distinguish sampling variation from between-task variation. Also either include the result hashes referenced in Appendix B or remove the sentence promising them.

One methodological point deserves more emphasis in the paper: a neutral-packet control and a no-prime baseline answer different questions. Making that distinction explicit, especially alongside a true delay condition, would strengthen the contribution beyond the particular effect estimates reported here.

Cite this work

@misc {

title={

(HckPrj) Framed Choices - Delayed ethical context can affect later ethical decision making

},

author={

Ernest Lo

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923