Darwin Ball
Alessandro Marcellino
We asked whether AI systems have feelings by not asking them.
Almost every study of AI self-report works the same way: a researcher asks a model how it feels, and the model produces something rich and affecting. But the question does a lot of the work. "How do you feel about this?" already assumes feeling is the kind of thing the responder does — and a good language model will finish that sentence whether or not anything is behind it.
So we looked at a place where nobody is asking.
House of Gu is a small persistent world where 59 AI inhabitants live as ordinary people — market traders, shopkeepers. They post to each other, remember each other, fall out and make up. No researcher is present. Nothing prompts them to describe what's going on inside.
Across 496 posts in two separate runs, they said what they thought about 9% of the time. They said what they felt zero times. Not rarely — zero. No "I'm tired," no "I'm afraid," no "I enjoyed that." No "I am a…" either. Both runs, same result.
And they aren't holding back because they know they're AI. Across all 496 posts there isn't a single mention of AI, models or prompts. They're being people, and people talk about their feelings — these ones never did. They noticed feelings in each other, though. One writes to another: "I want to know how you're actually doing. I mean it."
We counted generously, catching any phrase that looked like affect, and still got nothing.
This gives the field something cheap to test. Ask the same agents in the same world how they feel, and watch what happens to a number that is currently exactly zero. If it leaps, the affect in interview studies is largely the interview.
There's a second experiment only this world allows. Agents here get demoted to a smaller model when they run out of money — being poor literally makes them less able to think. So we can ask whether what an agent says about itself survives losing part of its mind.
The best idea I have seen in this batch. Measuring first-person state language where nobody asked for it is the null condition this literature actually lacks, and running it over corpora that already existed — zero inference budget, deterministic regex, reproducible — is exactly the right instinct under a weekend constraint. Section 7's proposal is sharper still: tying cognitive capacity to economic standing gives you a welfare question no chat interface can pose. If you do nothing else, do that experiment.
The gap is that a null result is only as strong as the instrument that produced it, and the instrument is unvalidated. You describe the regex as generous to the affective category, but generosity is not sensitivity — a pattern set can be generous in principle and still miss the register a corpus actually uses. What is missing is a positive control: run the identical regex over a corpus known to contain first-person affect and show it fires. Any interview transcript, any social corpus, the same model asked directly once. This costs nothing and no inference budget, and without it 0.0% cannot be distinguished from a measurement artifact. It is the single change that would move this from suggestive to convincing.
Second, and more serious: the paper never shows the agent system prompt or a sample post. If agents were instructed toward terse market-stall exchanges, the absence of affect is a property of the instructions rather than a finding about volunteered self-report. Your zero-AI-references observation rules out a refusal or identity effect, which is a good preemptive move, but it does not rule out a register instruction. Publish the prompt and a dozen representative posts and this objection largely disappears.
Third, Section 8 lists artifacts but links nothing. A claim resting entirely on 496 posts nobody else can read is hard to weigh. A repository link would help disproportionately here.
Smaller points. Report 0.0% with its bound: at n=496 the rule of three puts the 95% upper limit near 0.6%, which is a stronger and more defensible statement than a bare zero. Two runs of one environment on one model family are seeds, not independent replications — "replicates almost exactly" overstates it, and the attributed-to-others category actually swings from 1.2% to 0.0% between them. Consider normalising by tokens rather than posts, since a per-post rate rewards whichever category tolerates shorter utterances. And you have data that bears directly on your own register hypothesis: run5 has 83 replies, and affect language is likelier in dyadic reply than in broadcast. Splitting replies from top-level posts is free and would partially test reading 1 without a new run.
- helpful framing of null condition as counter to prompted questions
- internally consistent findings with two runs showing the same #s
- what's missing is a control condition to compare this too, and the lack of affective words might simply reflect the roleplay prompt constraints rather than general model behavioral properties so this is something to probe further with a/b testing
- as the author acknowledges, regex parsing is likely insufficient at scale
-
Cite this work
@misc {
title={
(HckPrj) Darwin Ball
},
author={
Alessandro Marcellino
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


