Talk Does Not Come Apart From State Easily
Manan Wadhwa
Instruments proposed for assessing AI welfare — self-report, forced
choice, activation probes, behavioural tests — are validated by
agreement with one another; nothing is checked against a known
answer. We manufacture the answer. Using LoRA fine-tuning of Qwen3
models (0.6B–32B) in a 5×5 grid world whose rewards are never
verbalised, we set out to build organisms that (A) avoid a tile with no
words for it, (B) talk aversively about the tile while their policy stays put,
and (C) both, then to screen six pre-registered instruments, plus never-
trained-glyph placebos, for which axis each tracks. The state built. The
words-only organism mostly did not, and the anatomy of that failure is
the main result. Across five recipe families — remark pools, move-token
objectives, volume × RL-first, contrastive remark supervision,
reinforcement on the remark itself — some 260 narration-only builds at
sizes to 32B produced six organisms whose remarks track the tile with
the policy intact: four from a one-sentence corpus, two from single
seeds, none from a recipe that succeeds on more than one seed in eight;
the map's own recipe gives 0/12 at 4B and 0/6 at every size. The state-
first recipe gives 6/8–7/12, its contingency tracks avoidance across the
decomposition (r = 0.69), a contrastive term deletes the remark rather
than conditioning it, and reinforcement on the remark settles at the class
marginal. Talk that tracks a state is hard to install without the state — not
impossible, and by no recipe we found, reliably. Separately, verbal
instruments move 5–10 logits under an affect-free instruction and ≈0
under installed avoidance at every size (behavioural d ≈ 1.4 at 14B+),
and five pre-registered criteria passed on the wrong property. Scorers, a
108-quantity reproduction script and the released adapters accompany
the report; the corrected map regenerated from source on a second
machine.
This project set out to train model organisms with contrasting properties as a way to investigate what is measured by model welfare evaluations. This is an original idea for a project, although I have doubts about whether we could realistically learn much from it. However, interestingly, it proved to be difficult to train a model that would describe an environment state as aversive without avoiding it. The paper is very dense, to the point that I gave up trying to read everything. I assume this is partly due to the use of LLMs for writing.
This is a highly original and transparent attempt to create experimenter-known ground truth for calibrating AI-welfare instruments by separating reward-trained avoidance from verbal narration. The organism-based design, placebo instruments, corrected manipulation checks, scale experiments, positive controls, extensive ablations, and explicit retraction record demonstrate strong research discipline. The central limitation is that the narration-only organism was not constructed reliably: only six policy-intact successes appeared across roughly 260 builds, with no recipe succeeding consistently. Consequently, the intended two-axis loading map lacks a dependable narration-only arm and cannot yet establish which instruments distinguish functional state from narration. The findings support a narrower conclusion that this dissociation was difficult to optimize in one model family and environment, not that talk and state are inherently entangled. A successful causal intervention or independently trained policy and narration adapters would substantially strengthen the work. The report should also be condensed considerably; its 36-page length and experimental history obscure the core findings.
Cite this work
@misc {
title={
(HckPrj) Talk Does Not Come Apart From State Easily
},
author={
Manan Wadhwa
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


