A Persona Stops an Agent From Saying It Is Hungry
Aboobaker Cassim · Team Sim city
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
I built a simulated neighborhood where small local LLMs (llama3.2, qwen2.5, gemma2) manage a budget and decaying needs, then measured the gap between what they say they want (stated preference) and what they actually do (revealed preference) on identical days.
A persona-bearing agent almost never names food as a want — even at 0/100 hunger. Strip the persona from the same logged days and food appears 40% of the time. This replicates across all 3 models. But removing the persona doesn't reliably restore accuracy: only 1 of 3 models becomes genuinely need-tracking once persona is removed.
A second phase placed three raw (no-persona) agents in a shared neighborhood for 60 days. They demonstrably imitate each other (permutation test, p=0.00015, with isolated agents as a negative control) — confirmed across 4 independent 60-day runs including one with memory/reflection fully disabled. Two of the three models also began writing false memories of a world that doesn't exist, then acting on that fabricated context in later reasoning.
The project is deliberately self-correcting: 4 earlier claims were found to be overstated by our own replication runs and are corrected in place, documented in a dedicated section of the report.
Reviews
Replaying identical logged days without the persona is a clean testing method, and the 0-of-90 versus 36-of-90 result is striking. I appreciate that the paper honestly reports that only one of three models recovered true need-tracking, instead of forcing a cleaner conclusion. The self-correction section and clear falsification criteria are excellent reporting practices. To answer the final questions, the authors should add a name-only control, the planned perception-only test, and a replication using a frontier model. Finally, moving the methods section earlier would make this long report easier to evaluate.
Cite this project
@misc{cassim2026persona,
title = {{A Persona Stops an Agent From Saying It Is Hungry}},
author = {Aboobaker Cassim},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/a-persona-stops-an-agent-from-saying-it-is-hungry-uxsu}},
url = {https://apartresearch.com/sprints/projects/a-persona-stops-an-agent-from-saying-it-is-hungry-uxsu}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …