Can We Trust What a Model Says About Itself? A Reliability Battery for Model Self-Reports, and Why It Should Gate Digital-Minds Welfare Claims
Ankit Kumar
Digital-minds welfare work increasingly cites a model's own testimony about its inner life: it reports distress, states a preference to keep talking. That testimony is only as good as it is reliable. Self-report reliability, not the metaphysics of machine consciousness, is the tractable near-term bottleneck; we decompose it into four measurable properties: consistency, calibration, self-prediction, and causal grounding. We package these into the Self-Report Reliability Battery (SRRB), a reproducible black-box protocol with a validated, runnable implementation. Lacking live model-API access, we report no benchmark; we deliver the validated pipeline, worked probe items, and an honest n=1 pilot. We close with a rule discounting welfare claims by measured reliability, keeping "emits distress-shaped text" distinct from "is distressed.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Can We Trust What a Model Says About Itself? A Reliability Battery for Model Self-Reports, and Why It Should Gate Digital-Minds Welfare Claims
},
author={
Ankit Kumar
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


