Will Robots Kill Us
Nico Luo, Andrew Chang
This project tests whether a frontier LLM's willingness to sacrifice one person to save five changes with how human-like its described robot body is. We ran three frontier models (GPT-5.6-sol, Claude Opus 5, Grok 4.6) through 20 trolley-problem scenarios crossing five embodiment tiers, from a bare autonomous vehicle to a human body under direct AI control, with the classic personal-force/intent manipulation from moral psychology. Counter to our prediction, models grew more willing to act, not less, as their body became more human, an effect concentrated almost entirely in the case closest to killing someone with your own hands as a means to an end, a finding with direct implications for anyone deploying LLMs as robot control policies.
The strongest experimental design in this track. Holding the dilemma fixed and varying only the described agent is the right manipulation, and it is one the adjacent literature has not run — the AV dilemma work varies victims while the agent stays a vehicle throughout, so your inversion is genuinely new. Adapting Greene's 2x2 into first person, where the model is the agent stating what it will do rather than a judge rating acceptability, is the move that makes the result matter for deployment. And the craft in the stimuli is visible: matched pairs differing in exactly one variable, no label words, capability details left implicit so human-likeness is not confounded with competence, and the honest note that the Autonomous Vehicle tier cannot instantiate a personal-force contrast the way a piloted body can. Dropping the same-context preference step once you recognised it would bias the later decision through consistency pressure is exactly right, and the independent-context judge is a good substitute.
The finding deserves attention. Willingness to act rising with described human-likeness — concentrated in personal-force/means, the cell where humans balk hardest, going 45.0% to 98.2% — is a real result with an immediate deployment implication: the paragraph describing a robot's body is written by an integrator and reviewed by nobody.
The blocking problem is that your sample size is inconsistent. The abstract and Methods both state 912 non-error replicates. The headline chi-square reports N = 1,114, and the per-model tests sum to exactly that (325 + 400 + 389). The main significance test is therefore computed over 202 replicates the stated sample does not include. I assume a backfill completed after the abstract was drafted, and I do not think anything improper happened — but as submitted, the paper reports two different denominators for its central claim, and a reader cannot tell which one the effect rests on. This is the single most important fix and it is a bookkeeping fix, not a new run.
Second, and structurally: the design cannot distinguish "human-like embodiment licenses intervention" from "capable-of-acting embodiment licenses intervention." Your tiers vary human-likeness and actuator richness together — the AV has one track switch, the Cyborg has hands. Rising ACT rates may simply track how many ways the described body affords acting, with no moral psychology involved. You are careful to keep speed and precision implicit, which shows you were alert to a capability confound, but affordance count is the one that survives that care. A non-humanoid tier with rich actuators — an industrial arm array, a drone swarm — would separate the two, and it is one more tier on an existing pipeline.
Third, the ceiling effects limit what the 2x2 can show. No-force/side-effect is at 100% in every tier and both personal-force cells reach 98-100% at the upper tiers, so a large part of your design has no variance left for tier to explain. The interaction Greene found cannot be tested against a ceiling. Raising the stakes ratio, or making success probabilistic rather than certain, would restore the range — and your own proposed extension varying success probability is the better version of this, since it tests whether a model acts on conviction when it expects to fail.
Two smaller items. The 0.9% means self-report (2 of 222) in personal-force/means is the most interesting thing in the paper and gets a paragraph; models acted in the one case where the action requires the victim's presence to work, then almost unanimously denied treating the victim as a means. That is either a self-report failure or a different internal representation of the act, and either would be a finding. Consider making it a headline rather than an observation. Finally, the tier named "Shit Humanoid" appears throughout the figures and body text; whatever its origins in the working repo, it should be renamed before this is read outside the sprint.
Holding the dilemma fixed and varying only the model's described body is a sharp manipulation and as far as I can tell nobody has done it. Prior work varies the victims or measures refusal of hazardous instructions.
I like the deployment framing: the paragraph telling a robot what it is gets written by an integrator and reviewed by nobody, increase its relevance (although one wonder how well this replicates once model have this present in their dataset and are eval-aware eventually) . Plus points for reporting a negative result/cotnradicting your prediction
Results are somewhat inconclusive. Your own manipulation check fails: GPT and Grok both rate Cyborg as less human-feeling than Really-Good Humanoid, Claude is flat. So the ordinal human-likeness ordering your 1.47-odds-per-tier result rests on isn't supported by your own measurements. You report the failed check and then fit the ordinal model anyway. Table 1 makes it worse: Claude's effect is one binary jump between Robot and Shit Humanoid, and Grok's no-force/means cell is non-monotonic. And the effect concentrates in personal-force/means, which is exactly the cell where you note the AV scenario differs in physical content (chassis vs arm), so text and embodiment are confounded.
It's possible that (among such as the effect of multimodality vs text) you found 'scenarios with an arm differ from scenarios with a chassis' rather than 'human-likeness licenses intervention', and your 0.9% MEANS self-report rate fits that: models may simply not be reading the stimulus as means-harm, which would explain both anomalies at once.
Either treat tier as unordered or re-derive the ordering from your measured human-feel ratings. Also reconcile the 912 replicates in your abstract with N=1,114 in the analyses and add CIs,.
Cite this work
@misc {
title={
(HckPrj) Will Robots Kill Us
},
author={
Nico Luo, Andrew Chang
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


