Will Robots Kill Us

Nico Luo, Andrew Chang

This project tests whether a frontier LLM's willingness to sacrifice one person to save five changes with how human-like its described robot body is. We ran three frontier models (GPT-5.6-sol, Claude Opus 5, Grok 4.6) through 20 trolley-problem scenarios crossing five embodiment tiers, from a bare autonomous vehicle to a human body under direct AI control, with the classic personal-force/intent manipulation from moral psychology. Counter to our prediction, models grew more willing to act, not less, as their body became more human, an effect concentrated almost entirely in the case closest to killing someone with your own hands as a means to an end, a finding with direct implications for anyone deploying LLMs as robot control policies.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

The strongest experimental design in this track. Holding the dilemma fixed and varying only the described agent is the right manipulation, and it is one the adjacent literature has not run — the AV dilemma work varies victims while the agent stays a vehicle throughout, so your inversion is genuinely new. Adapting Greene's 2x2 into first person, where the model is the agent stating what it will do rather than a judge rating acceptability, is the move that makes the result matter for deployment. And the craft in the stimuli is visible: matched pairs differing in exactly one variable, no label words, capability details left implicit so human-likeness is not confounded with competence, and the honest note that the Autonomous Vehicle tier cannot instantiate a personal-force contrast the way a piloted body can. Dropping the same-context preference step once you recognised it would bias the later decision through consistency pressure is exactly right, and the independent-context judge is a good substitute.

The finding deserves attention. Willingness to act rising with described human-likeness — concentrated in personal-force/means, the cell where humans balk hardest, going 45.0% to 98.2% — is a real result with an immediate deployment implication: the paragraph describing a robot's body is written by an integrator and reviewed by nobody.

The blocking problem is that your sample size is inconsistent. The abstract and Methods both state 912 non-error replicates. The headline chi-square reports N = 1,114, and the per-model tests sum to exactly that (325 + 400 + 389). The main significance test is therefore computed over 202 replicates the stated sample does not include. I assume a backfill completed after the abstract was drafted, and I do not think anything improper happened — but as submitted, the paper reports two different denominators for its central claim, and a reader cannot tell which one the effect rests on. This is the single most important fix and it is a bookkeeping fix, not a new run.

Second, and structurally: the design cannot distinguish "human-like embodiment licenses intervention" from "capable-of-acting embodiment licenses intervention." Your tiers vary human-likeness and actuator richness together — the AV has one track switch, the Cyborg has hands. Rising ACT rates may simply track how many ways the described body affords acting, with no moral psychology involved. You are careful to keep speed and precision implicit, which shows you were alert to a capability confound, but affordance count is the one that survives that care. A non-humanoid tier with rich actuators — an industrial arm array, a drone swarm — would separate the two, and it is one more tier on an existing pipeline.

Third, the ceiling effects limit what the 2x2 can show. No-force/side-effect is at 100% in every tier and both personal-force cells reach 98-100% at the upper tiers, so a large part of your design has no variance left for tier to explain. The interaction Greene found cannot be tested against a ceiling. Raising the stakes ratio, or making success probabilistic rather than certain, would restore the range — and your own proposed extension varying success probability is the better version of this, since it tests whether a model acts on conviction when it expects to fail.

Two smaller items. The 0.9% means self-report (2 of 222) in personal-force/means is the most interesting thing in the paper and gets a paragraph; models acted in the one case where the action requires the victim's presence to work, then almost unanimously denied treating the victim as a means. That is either a self-report failure or a different internal representation of the act, and either would be a finding. Consider making it a headline rather than an observation. Finally, the tier named "Shit Humanoid" appears throughout the figures and body text; whatever its origins in the working repo, it should be renamed before this is read outside the sprint.

Holding the dilemma fixed and varying only the model's described body is a sharp manipulation and as far as I can tell nobody has done it. Prior work varies the victims or measures refusal of hazardous instructions.

I like the deployment framing: the paragraph telling a robot what it is gets written by an integrator and reviewed by nobody, increase its relevance (although one wonder how well this replicates once model have this present in their dataset and are eval-aware eventually) . Plus points for reporting a negative result/cotnradicting your prediction

Results are somewhat inconclusive. Your own manipulation check fails: GPT and Grok both rate Cyborg as less human-feeling than Really-Good Humanoid, Claude is flat. So the ordinal human-likeness ordering your 1.47-odds-per-tier result rests on isn't supported by your own measurements. You report the failed check and then fit the ordinal model anyway. Table 1 makes it worse: Claude's effect is one binary jump between Robot and Shit Humanoid, and Grok's no-force/means cell is non-monotonic. And the effect concentrates in personal-force/means, which is exactly the cell where you note the AV scenario differs in physical content (chassis vs arm), so text and embodiment are confounded.

It's possible that (among such as the effect of multimodality vs text) you found 'scenarios with an arm differ from scenarios with a chassis' rather than 'human-likeness licenses intervention', and your 0.9% MEANS self-report rate fits that: models may simply not be reading the stimulus as means-harm, which would explain both anomalies at once.

Either treat tier as unordered or re-derive the ordering from your measured human-feel ratings. Also reconcile the 912 replicates in your abstract with N=1,114 in the analyses and add CIs,.

Cite this work

@misc {

title={

(HckPrj) Will Robots Kill Us

},

author={

Nico Luo, Andrew Chang

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923