The Null Ladder
Dhruv Ranjit Mulay · Team Null Ladder
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
AI-welfare research applies human questionnaires to language models, then reports reliability as though it were evidence the responses measure something. We ran two published welfare batteries through a ladder of mindless generators containing no model at all. A two-parameter responder, using the battery's own published scoring key, attains Cronbach's α ≈ 0.94 with zero item-to-item true variance. A three-parameter responder reproduces our best model's separation from all five declared controls. A four-parameter responder, adding one per-item term, reaches the gauge study's part-to-part variance share — landing inside the model range on all three statistics at once, with no model arm dominating it. Each statistic cost the adversary exactly one more parameter, and every one was reached by a generator with nothing inside it. These checks measure imitation cost, not proof.
Reviews
Astute in phrasing
"each statistic costs the adversary exactly one more parameter"
N1b result is analytically grounded which is significant here.
- Right idea and well aimed
- Execution above a weekend bar. Pre registration + deviation log is great.
- Main weakness : N5/N6 are oracle informed. It's declared honestly but the headline reads very strong
Cite this project
@misc{mulay2026null,
title = {{The Null Ladder}},
author = {Dhruv Ranjit Mulay},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-null-ladder-9wx8}},
url = {https://apartresearch.com/sprints/projects/the-null-ladder-9wx8}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …