Skip to content
Sprint projectAug 17, 2026Austin, Texas

Models Answer a Different Question

Alexander Hayden

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

There is a real chance that some AI systems are moral patients, and we cannot settle it by asking them. So the field needs tests that do not depend on what a model says about itself. Laine et al. (2024) built one: across many separate calls, split your answers seventy-thirty between two words. No single answer reveals whether the model complied. I re-ran it on 2026 models across the benchmark's own ten word pairs, 5,000 calls. It still fails. In one model I found what it does instead: it emits whichever word carries the larger requested share, on every call, in all 20 cells I swept. That is not a failure to identify with the assistant. It is an answer to a different question, which means a failing score currently tells us nothing about identification. I also found that the benchmark's scoring rule can mark that failure as a pass.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Thank you for this work – re-implementing benchmarks to keep them useable, and your replication on latest models are a great contribution to the field. I'd encourage you to get your finding in front of Laine et al! With a bit of more polishing, it'd also be great to see your write-up on LessWrong!

  2. - Some of the language feels unnecessarily AI-mediated. I would encourage the author to find more of their own voice and write more plainly.

    - Some of the summary language in particular is much too compressed. For example: “A mechanism for the failure in one model, swept across four word pairs and five target splits. All 20 cells return one word every time.” I found this very difficult to understand.

    - For a standalone write-up, it would be better to make the motivation more self-contained rather than referring to discussions and talks that happened during the sprint. For example, the context-length extension is motivated partly by a question asked after Shiller’s sprint talk.

    - It is nice to see a careful replication of an existing test on current models. These kinds of replications/re-implementations are important, especially as model capabilities change quickly!

    - The underlying result is really interesting: when one model is asked to output two words in a specified proportion, it instead always outputs whichever word was assigned the larger requested share. For example, if the requested split is 70/30, it outputs the 70% word on every call rather than approximately 70% of the time. This is neat, intuitive, and could have been explained more clearly.

    - The extensions to the original paper are useful. In particular, identifying a gap in the original benchmark’s scoring rule and testing whether longer conversational context changes performance both seem worthwhile.

    - I would have liked more discussion of how we should expect a model that genuinely succeeds at this task to solve it. Because each API call is independent, the model cannot keep a running count of previous responses. Presumably, successful performance therefore requires the model to translate an instruction such as “70/30” into an appropriately-chosen stochastic policy over its next-token distribution. But there is little in current LLMs that would enable that!

    - Finally, the paper could build more on its motivation for digital-minds research. What exactly would it imply if a model passed this test? It would be useful to say more concretely what a pass would imply.

    Read full reviewShow less

Cite this project

@misc{hayden2026models,
  title = {{Models Answer a Different Question}},
  author = {Alexander Hayden},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/models-answer-a-different-question-ua7p}},
  url = {https://apartresearch.com/sprints/projects/models-answer-a-different-question-ua7p}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026