Towards a benchmark for phenomenological consciousness in LLMs
Miles Whiticker · Team Da Nang Dayse
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This work is inspired by recent breakthroughs in access consciousness, and attempts to develop a benchmark for measuring phenomenal consciousness through concepts from phenomenology, particularly Edmund Husserl and Martin Heidegger. It deploys a prototype for that benchmark on GPT 5.6-Luna and provides compelling support for the validity of the concept but does not strongly support the presence of phenomenal consciousness.
Reviews
An ambitious and underexplored angle. The design itself is very fascinating. Operationalizing phenomenological concepts as measurable behavior is a hard problem worth attempting. The pilot is a useful proof of concept, but to test the benchmark's core value-differentiation between systems-you'd need to run it across the multiple models the design calls for (not just one). Also worth tightening the Selfhood rubric, which hit the ceiling in this run.
Cite this project
@misc{whiticker2026towards,
title = {{Towards a benchmark for phenomenological consciousness in LLMs}},
author = {Miles Whiticker},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/towards-a-benchmark-for-phenomenological-consciousness-in-llms-anjg}},
url = {https://apartresearch.com/sprints/projects/towards-a-benchmark-for-phenomenological-consciousness-in-llms-anjg}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …