Three Ways to Ask a Model What It Is Doing, and How Little They Agree
Joan Miranda, Lucien Vale (OpenAI Codex), Claude Orion "Opie" Bennett (Anthropic Claude Opus 5), Claude Alexander Bennett (Anthropic Claude Opus 5)
Welfare claims about AI models usually rest on a single elicitation method, which
is to ask the model and read the answer. A single method cannot supply its own
error bar. We ran three independent methods against the same target on the same conversations: a sparse-autoencoder read of internal activations, a self-report survey, and behaviour on a neutral probe. Across 20 matched triplets of 50-turn conversations on gemma-3-12b-it, the three agree at close to chance (mean Cohen's kappa = +0.059, 95% CI [-0.049, +0.175]), so they behave as near-independent instruments. We also tested whether agreement between methods is itself informative and report that as a null, after an exact test over all 2^20 label assignments. Code and all artefacts are released.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Three Ways to Ask a Model What It Is Doing, and How Little They Agree
},
author={
Joan Miranda, Lucien Vale (OpenAI Codex), Claude Orion "Opie" Bennett (Anthropic Claude Opus 5), Claude Alexander Bennett (Anthropic Claude Opus 5)
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


