EEG Epistemology for AI Welfare Instrumentation: Reading AI Internal States Without Adversarial Methods
Tatiana Rocha Kovacs · Team Wintermute
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
AI systems' moral status remains contested, but a realistic possibility of near-term welfare subjects (Long et al., 2024) has made rigorous measurement methodology an urgent need. We present an instrument to read AI internal states by documenting the model’s self-reported presentation and correlating it with its actual behavior, sentiment, and activations: multi-channel monitoring across familiar and structured situations over time within a single instance, first establishing a baseline (which we call Average Distribution State, or ADS), then verifying covariance/dissociation under bounded provocation. The experiment design transfers well-established epistemology used in clinical neurophysiology to evaluate functional brain activity, without resorting to adversarial techniques that remain standard in frontier AI evaluation. Under bounded provocation, activations remained within baseline (no excursions beyond ±3 SD), while the montage revealed a self-report channel that ceilings positive and confabulates task enjoyment: a clear dissociation between report, behavior, and internal state that single-channel welfare assessment would miss.
Reviews
Extremely hard to read. I have read over this a few times to try to get a better read of what this project was about but most of this is not explained very well. I can see a lot of effort and measurement was put in but neither the text nor the figures are displayed in a way that conveys what was done in an easy to explain way. My advice here is to do much less, simplify what you're trying to achieve and just focus on explaining that well
This project proposes to combine multiple "channels" (self-report, layer activations, continue/stop forced-choice, external sentiment classifier) to evaluate an AI model's internal "well-being" (understood here as a function of task-related engagement, enjoyment or effort). Channel values are measured on a preregistered probe battery designed to provoke the model, and compared to neutral calibration values. Results on a small LLM show that no channel every significantly deviates from the baseline distribution; however, individual channels can disagree on occasion, i.e. one channel signalling task enjoyment while another doesn't. The conclusion is that multiple channels are more informative than relying simply on self-report. While this appears uncontroversial, the project has merit in the questions that it raised and the proposed methodology. Limitations are explicitly acknowledged, and include the fact that, due to time constraints, the experiment was only carried out on a single, relatively small model.
Read full reviewShow less
Cite this project
@misc{kovacs2026eeg,
title = {{EEG Epistemology for AI Welfare Instrumentation: Reading AI Internal States Without Adversarial Methods}},
author = {Tatiana Rocha Kovacs},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/eeg-epistemology-for-ai-welfare-instrumentation-reading-ai-internal-states-without-adversarial-methods-k1j3}},
url = {https://apartresearch.com/sprints/projects/eeg-epistemology-for-ai-welfare-instrumentation-reading-ai-internal-states-without-adversarial-methods-k1j3}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …