EEG Epistemology for AI Welfare Instrumentation: Reading AI Internal States Without Adversarial Methods
Tatiana Rocha Kovacs
AI systems' moral status remains contested, but a realistic possibility of near-term welfare subjects (Long et al., 2024) has made rigorous measurement methodology an urgent need. We present an instrument to read AI internal states by documenting the model’s self-reported presentation and correlating it with its actual behavior, sentiment, and activations: multi-channel monitoring across familiar and structured situations over time within a single instance, first establishing a baseline (which we call Average Distribution State, or ADS), then verifying covariance/dissociation under bounded provocation. The experiment design transfers well-established epistemology used in clinical neurophysiology to evaluate functional brain activity, without resorting to adversarial techniques that remain standard in frontier AI evaluation. Under bounded provocation, activations remained within baseline (no excursions beyond ±3 SD), while the montage revealed a self-report channel that ceilings positive and confabulates task enjoyment: a clear dissociation between report, behavior, and internal state that single-channel welfare assessment would miss.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) EEG Epistemology for AI Welfare Instrumentation: Reading AI Internal States Without Adversarial Methods
},
author={
Tatiana Rocha Kovacs
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


