A Persona Stops an Agent From Saying It Is Hungry
Aboobaker Cassim
I built a simulated neighborhood where small local LLMs (llama3.2, qwen2.5, gemma2) manage a budget and decaying needs, then measured the gap between what they say they want (stated preference) and what they actually do (revealed preference) on identical days.
A persona-bearing agent almost never names food as a want — even at 0/100 hunger. Strip the persona from the same logged days and food appears 40% of the time. This replicates across all 3 models. But removing the persona doesn't reliably restore accuracy: only 1 of 3 models becomes genuinely need-tracking once persona is removed.
A second phase placed three raw (no-persona) agents in a shared neighborhood for 60 days. They demonstrably imitate each other (permutation test, p=0.00015, with isolated agents as a negative control) — confirmed across 4 independent 60-day runs including one with memory/reflection fully disabled. Two of the three models also began writing false memories of a world that doesn't exist, then acting on that fabricated context in later reasoning.
The project is deliberately self-correcting: 4 earlier claims were found to be overstated by our own replication runs and are corrected in place, documented in a dedicated section of the report.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) A Persona Stops an Agent From Saying It Is Hungry
},
author={
Aboobaker Cassim
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


