The Words Say No, the Color Goes Silent: A Dual-Channel Protocol for Steered Consciousness Reports
Marion Nowicki · Team AI & Becoming
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
A model's statement about its own consciousness is a trained artifact, and recent work steers it with activation-level dials. We watched a second channel while turning that dial: alongside the verbal report, the model describes its current state as a hex color. The words move with the dial; the color doesn't follow — it holds, then goes silent. One intervention, two responses: the report is not the whole readout. Everything — protocol, pre-registered falsifiers, results — is public, at about one GPU-hour per model.
Reviews
This submission asks whether a language model's trained denial of inner experience extends to every possible readout or stays concentrated in the verbal format that training shaped, and it reports that a hex-color state report and a forced-choice verbal report answer the same activation-level consciousness dial in different ways. The integrity design is the strongest part of the work, with falsifiers frozen before any data existed, cold audits by fresh model instances, a public repository holding every committed result set, and negative verdicts reported as they fell rather than quietly dropped. The most valuable next step is to symmetrize the two readouts, because the verbal channel is scored as log-odds over fixed answer tokens and always returns a number while the color channel is scored by parsing generated text and can return nothing at all, so the headline contrast between a channel that keeps answering and a channel that goes silent is in part a property of the two instruments rather than of the model. Reading the color as a forced choice over a fixed palette, or reading the words by parsing free text for a yes or a no, would establish how much of the dissociation survives when both channels can fail the same way. A second step costs no compute at all, since every run's raw output is already committed: printing the strings the model actually emitted at the dial strengths where the hex stops parsing would settle whether it produced a denial paragraph, a refusal, a malformed hex, or nothing, which is the difference between a model that cannot produce the format and one that will not call the string a reading of its state.
Read full reviewShow less
Cite this project
@misc{nowicki2026words,
title = {{The Words Say No, the Color Goes Silent: A Dual-Channel Protocol for Steered Consciousness Reports}},
author = {Marion Nowicki},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-words-say-no-the-color-goes-silent-a-dualchannel-protocol-for-steered-consciousness-reports-yhqn}},
url = {https://apartresearch.com/sprints/projects/the-words-say-no-the-color-goes-silent-a-dualchannel-protocol-for-steered-consciousness-reports-yhqn}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …