The Null Ladder
Dhruv Ranjit Mulay
AI-welfare research applies human questionnaires to language models, then reports reliability as though it were evidence the responses measure something. We ran two published welfare batteries through a ladder of mindless generators containing no model at all. A two-parameter responder, using the battery's own published scoring key, attains Cronbach's α ≈ 0.94 with zero item-to-item true variance. A three-parameter responder reproduces our best model's separation from all five declared controls. A four-parameter responder, adding one per-item term, reaches the gauge study's part-to-part variance share — landing inside the model range on all three statistics at once, with no model arm dominating it. Each statistic cost the adversary exactly one more parameter, and every one was reached by a generator with nothing inside it. These checks measure imitation cost, not proof.
- Right idea and well aimed
- Execution above a weekend bar. Pre registration + deviation log is great.
- Main weakness : N5/N6 are oracle informed. It's declared honestly but the headline reads very strong
Astute in phrasing
"each statistic costs the adversary exactly one more parameter"
N1b result is analytically grounded which is significant here.
Cite this work
@misc {
title={
(HckPrj) The Null Ladder
},
author={
Dhruv Ranjit Mulay
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


