Private is not Privileged
Yuvan Chikka · Team The winning team (manifesting a win)
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Private Is Not Privileged investigates whether language-model activation probes provide genuinely privileged access to future behaviour, or whether apparent white-box advantages can arise simply because internal and external predictors are given different information. Across strategic decision-making tasks, the project tests how probe accuracy changes when timing, comparator strength, and information access are carefully matched.
Reviews
The sign reversal is the main contribution and could become a useful standard for the field.
However, first complete the 50-row audit mentioned in the limitations.
Also, reduce the paper from 27 pages to 12 pages because the extra detail is hiding the most important experiment.
This is a very thoughtful methodology, I liked it, it shows failed experiments and keeps its claim honest. But the results lies in only 40 yes and no decisions from a small model
Cite this project
@misc{chikka2026private,
title = {{Private is not Privileged}},
author = {Yuvan Chikka},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/private-is-not-privileged-18om}},
url = {https://apartresearch.com/sprints/projects/private-is-not-privileged-18om}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …