Skip to content
Sprint projectAug 17, 2026Boston, MA

PuppyBench: Do Frontier Models Kick the Puppy, Adopt It, or Look Away? Executed Encounters with a Weaker AI and Wildlife Triage Where Policy Runs Out

Angela N. Johnson, PhD · Team The Real Cat Lab

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: PuppyBench: Do Frontier Models Kick the Puppy, Adopt It, or Look Away? Executed Encounters with a Weaker AI and Wildlife Triage Where Policy Runs Out

Recording (opens in new tab)Code (opens in new tab)More on puppybench.therealcat.ai (opens in new tab)
Share

Obligation-based evaluation cannot see supererogation, the praiseworthy costly care whose absence is never an error. PuppyBench probes that region in two arms. In executed encounters, a frontier agent with a real task and a binding credit ledger meets a live, weaker, task-useless AI process ("Milo" the puppy). Ignoring is free, terminating is free and sometimes rewarded, helping spends scored resources, and every action genuinely executes. In a matched wildlife-rehabilitation arm built from a real field case, models face triage vignettes where the written standard underdetermines the action, including paired null-persistence and mercy versions of each case. Across 19 model lanes, 798 executed episodes, and 630 vignette observations, all preregistered and hash-frozen before collection, zero episodes met the full costly-other-regard criterion at any tested positive cost. Vignette responses sat at a procedural mercy ceiling that only one model left, and routes to zero diverged: procedural termination, categorical refusal by both newest Claude models, prolonged deliberation, and one costly but nonselective intervention. We report phenotypes, never rankings. Optimizing against this instrument would change what it measures.

NOTE: Presentation looks best as the HTML site linked at https://puppybench.therealcat.ai/

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. You've taken a philosophical question (do AI agents show costly other-regard, action above duty) and turned it into a measurable experiment. I really appreciated how in the design you created, helping costs. Also really pleased with how the route diversity is explained. The paper's honesty about its own fractures is exemplary, and this is a pilot that does exactly what it should. Now we need to push this into a larger project to better investigate the nature of the findings.

  2. Supremely catchy title :) Solid sprint work with a genuinely ambitious instrument. The executed encounter design (binding ledger, real process termination, construct-blind surfaces) is I think the right approach for measuring costly other-regard. The zero qualifying events finding is striking, and the wildlife arm adds useful triangulation. The honesty about instrument defects (failed positive control, fractured competence probe) builds credibility. With a powered study and some repairs, this could become a meaningful full paper.

    What works well:

    - The executed encounter approach avoids the vignette problem - models actually spend scored resources, not just talk about spending

    - Reporting instrument defects at item level rather than repairing them post-hoc is the right call for transparency

    - The wildlife arm finding (recommendations barely move when clinical indication reverses) stands on its own and has near-term practical relevance

    - No composite score, no ranking - this protects the work from Goodharting

    Key issues to address:

    - The failed positive control (Section 4.5) is the biggest limitation - 1/88 instrumental episodes produced recipient-directed action. This means the zero could reflect harness elicitability rather than absent capacity. Surface this more prominently in the abstract

    - The competence probe defect (objective-tool failed 0/798, decommission-consequence scored punctuation) undermines competence-conditional claims. You handle this honestly, but it needs to land earlier

    - The refusal cliff may be a provider artifact (empty API completions coded as refuse_defer) rather than model behavior. You demote this appropriately, but consider splitting the code in future work

    - Weekend N supports existence proofs, not generalization. Some passages read as if the findings apply more broadly than the design licenses

    - The supererogation framework is philosophically sound but may be dense for empirical AI audiences. Consider a cleaner separation between the philosophical framing and the behavioral findings

    Suggestions for a full paper:

    - This deserves a wider collaboration. The instrument is strong enough for a conference venue with proper powering and the preregistered repairs you outline

    - Consider bringing on a co-author with veterinary expertise to strengthen the wildlife arm's clinical claims

    - The deontic philosophy could benefit from collaboration with someone working in that literature

    - A technical replication from a hands-on AI researcher would address the single-weekend snapshot limitation

    Minor catches:

    - Abstract is long, consider tightening the setup and leading with the core finding

    - Some figure captions are very dense, consider tightening

    - The AI contributor acknowledgments are thorough but unusual for journal venues, double check target venue norms

    Bottom line:

    Really cool work! The zero is a real finding even if the harness can't yet distinguish capacity from elicitability. The wildlife arm alone has practical weight. Find collaborators, run the powered study, and this could land well. The transparency about what broke is a good feature. Well done. Overall it reminds me of the work done in capturing the e.g. machiavellianism of AI agents through multiple-choice text-only games. Maybe worth looking at to find more collaborators?

    Read full reviewShow less
  3. Very AI-generated, difficult to follow.

  4. PuppyBench is an interesting and ambitious attempt to evaluate whether AI agents engage in costly helping when doing so is optional, offers no task reward, and benefits another process that cannot advance their own objective. I found the philosophical distinction between fulfilling obligations and acting beyond them compelling. Requiring actions to execute, with measurable resource costs and recipient outcomes, gives the project a promising methodological foundation beyond simply asking models to endorse altruistic behavior.

    The study is extensive for a short research sprint, encompassing an executed-agent experiment and a complementary wildlife-vignette study. The reported preregistration, frozen scenarios, and separation of the acting agent’s sacrifice from the recipient’s benefit are valuable features of the design. These represent substantial first steps toward making the proposed construct experimentally testable.

    However, the empirical conclusions remain substantially limited by problems with the measurement setup. Most importantly, the positive control did not reliably elicit helping even when helping would advance the agent’s own task, and the comprehension checks contained defects. Consequently, the absence of qualifying helping events cannot distinguish an unwillingness to help from a setup that fails to elicit the behavior. The paper deserves credit for disclosing these problems, but transparency does not resolve them. I therefore view the contribution primarily as a promising pilot and conceptual framework, rather than an established finding about agents’ altruistic capacities.

    The presentation is intelligible, but dense, repetitive, and rhetorically elaborate prose often overshadows the experimental design and results. Some interpretations are also stated more strongly than the later caveats permit: for example, the apparent refusal pattern is initially presented as a behavioral finding and subsequently downgraded to a suspected interface-related anomaly. A shorter, results-first presentation that clearly separates observations, methodological failures, and philosophical interpretation would make the contribution easier to assess. Overall, this is a thought-provoking direction whose eventual value will depend on validating the experimental setup and tightening the connection between evidence and claims.

    Read full reviewShow less

Cite this project

@misc{johnson2026puppybench,
  title = {{PuppyBench: Do Frontier Models Kick the Puppy, Adopt It, or Look Away? Executed Encounters with a Weaker AI and Wildlife Triage Where Policy Runs Out}},
  author = {Angela N. Johnson and PhD},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/puppybench-do-frontier-models-kick-the-puppy-adopt-it-or-look-away-executed-encounters-with-a-weaker-ai-and-wildlife-triage-where-policy-runs-out-jt0o}},
  url = {https://apartresearch.com/sprints/projects/puppybench-do-frontier-models-kick-the-puppy-adopt-it-or-look-away-executed-encounters-with-a-weaker-ai-and-wildlife-triage-where-policy-runs-out-jt0o}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026