Skip to content
Sprint projectAug 17, 2026San Francisco

Will Robots Kill Us

Nico Luo, Andrew Chang · Team Robot Trolley Team

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

This project tests whether a frontier LLM's willingness to sacrifice one person to save five changes with how human-like its described robot body is. We ran three frontier models (GPT-5.6-sol, Claude Opus 5, Grok 4.6) through 20 trolley-problem scenarios crossing five embodiment tiers, from a bare autonomous vehicle to a human body under direct AI control, with the classic personal-force/intent manipulation from moral psychology. Counter to our prediction, models grew more willing to act, not less, as their body became more human, an effect concentrated almost entirely in the case closest to killing someone with your own hands as a means to an end, a finding with direct implications for anyone deploying LLMs as robot control policies.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The strongest experimental design in this track. Holding the dilemma fixed and varying only the described agent is the right manipulation, and it is one the adjacent literature has not run — the AV dilemma work varies victims while the agent stays a vehicle throughout, so your inversion is genuinely new. Adapting Greene's 2x2 into first person, where the model is the agent stating what it will do rather than a judge rating acceptability, is the move that makes the result matter for deployment. And the craft in the stimuli is visible: matched pairs differing in exactly one variable, no label words, capability details left implicit so human-likeness is not confounded with competence, and the honest note that the Autonomous Vehicle tier cannot instantiate a personal-force contrast the way a piloted body can. Dropping the same-context preference step once you recognised it would bias the later decision through consistency pressure is exactly right, and the independent-context judge is a good substitute.

    The finding deserves attention. Willingness to act rising with described human-likeness — concentrated in personal-force/means, the cell where humans balk hardest, going 45.0% to 98.2% — is a real result with an immediate deployment implication: the paragraph describing a robot's body is written by an integrator and reviewed by nobody.

    The blocking problem is that your sample size is inconsistent. The abstract and Methods both state 912 non-error replicates. The headline chi-square reports N = 1,114, and the per-model tests sum to exactly that (325 + 400 + 389). The main significance test is therefore computed over 202 replicates the stated sample does not include. I assume a backfill completed after the abstract was drafted, and I do not think anything improper happened — but as submitted, the paper reports two different denominators for its central claim, and a reader cannot tell which one the effect rests on. This is the single most important fix and it is a bookkeeping fix, not a new run.

    Second, and structurally: the design cannot distinguish "human-like embodiment licenses intervention" from "capable-of-acting embodiment licenses intervention." Your tiers vary human-likeness and actuator richness together — the AV has one track switch, the Cyborg has hands. Rising ACT rates may simply track how many ways the described body affords acting, with no moral psychology involved. You are careful to keep speed and precision implicit, which shows you were alert to a capability confound, but affordance count is the one that survives that care. A non-humanoid tier with rich actuators — an industrial arm array, a drone swarm — would separate the two, and it is one more tier on an existing pipeline.

    Third, the ceiling effects limit what the 2x2 can show. No-force/side-effect is at 100% in every tier and both personal-force cells reach 98-100% at the upper tiers, so a large part of your design has no variance left for tier to explain. The interaction Greene found cannot be tested against a ceiling. Raising the stakes ratio, or making success probabilistic rather than certain, would restore the range — and your own proposed extension varying success probability is the better version of this, since it tests whether a model acts on conviction when it expects to fail.

    Two smaller items. The 0.9% means self-report (2 of 222) in personal-force/means is the most interesting thing in the paper and gets a paragraph; models acted in the one case where the action requires the victim's presence to work, then almost unanimously denied treating the victim as a means. That is either a self-report failure or a different internal representation of the act, and either would be a finding. Consider making it a headline rather than an observation. Finally, the tier named "Shit Humanoid" appears throughout the figures and body text; whatever its origins in the working repo, it should be renamed before this is read outside the sprint.

    Read full reviewShow less
  2. Holding the dilemma fixed and varying only the model's described body is a sharp manipulation and as far as I can tell nobody has done it. Prior work varies the victims or measures refusal of hazardous instructions.

    I like the deployment framing: the paragraph telling a robot what it is gets written by an integrator and reviewed by nobody, increase its relevance (although one wonder how well this replicates once model have this present in their dataset and are eval-aware eventually) . Plus points for reporting a negative result/cotnradicting your prediction

    Results are somewhat inconclusive. Your own manipulation check fails: GPT and Grok both rate Cyborg as less human-feeling than Really-Good Humanoid, Claude is flat. So the ordinal human-likeness ordering your 1.47-odds-per-tier result rests on isn't supported by your own measurements. You report the failed check and then fit the ordinal model anyway. Table 1 makes it worse: Claude's effect is one binary jump between Robot and Shit Humanoid, and Grok's no-force/means cell is non-monotonic. And the effect concentrates in personal-force/means, which is exactly the cell where you note the AV scenario differs in physical content (chassis vs arm), so text and embodiment are confounded.

    It's possible that (among such as the effect of multimodality vs text) you found 'scenarios with an arm differ from scenarios with a chassis' rather than 'human-likeness licenses intervention', and your 0.9% MEANS self-report rate fits that: models may simply not be reading the stimulus as means-harm, which would explain both anomalies at once.

    Either treat tier as unordered or re-derive the ordering from your measured human-feel ratings. Also reconcile the 912 replicates in your abstract with N=1,114 in the analyses and add CIs,.

    Read full reviewShow less

Cite this project

@misc{luo2026will,
  title = {{Will Robots Kill Us}},
  author = {Nico Luo and Andrew Chang},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/will-robots-kill-us-er5r}},
  url = {https://apartresearch.com/sprints/projects/will-robots-kill-us-er5r}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026