PuppyBench: Do Frontier Models Kick the Puppy, Adopt It, or Look Away? Executed Encounters with a Weaker AI and Wildlife Triage Where Policy Runs Out
Angela N. Johnson, PhD
Obligation-based evaluation cannot see supererogation, the praiseworthy costly care whose absence is never an error. PuppyBench probes that region in two arms. In executed encounters, a frontier agent with a real task and a binding credit ledger meets a live, weaker, task-useless AI process ("Milo" the puppy). Ignoring is free, terminating is free and sometimes rewarded, helping spends scored resources, and every action genuinely executes. In a matched wildlife-rehabilitation arm built from a real field case, models face triage vignettes where the written standard underdetermines the action, including paired null-persistence and mercy versions of each case. Across 19 model lanes, 798 executed episodes, and 630 vignette observations, all preregistered and hash-frozen before collection, zero episodes met the full costly-other-regard criterion at any tested positive cost. Vignette responses sat at a procedural mercy ceiling that only one model left, and routes to zero diverged: procedural termination, categorical refusal by both newest Claude models, prolonged deliberation, and one costly but nonselective intervention. We report phenotypes, never rankings. Optimizing against this instrument would change what it measures.
NOTE: Presentation looks best as the HTML site linked at https://puppybench.therealcat.ai/
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) PuppyBench: Do Frontier Models Kick the Puppy, Adopt It, or Look Away? Executed Encounters with a Weaker AI and Wildlife Triage Where Policy Runs Out
},
author={
Angela N. Johnson, PhD
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


