A pre-registered REBUS analog fails in two instruct models for reasons its controls reveal
Akanksha Gupta · Team 3am Labs
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Model welfare evals often treat a 0–100 self-report as a state. I pre-registered a REBUS-shaped residual operator on Qwen2.5-3B-Instruct and Llama-3.2-3B-Instruct, with an accept rule locked before the run: diversity up, dissolution up ≥10, inflation within 5, and an entropy-matched temperature control that must not copy the rise. Both reject. Qwen’s dissolution score moves the wrong way. Llama’s write looks like a hit (+11) until T=1.6 reaches 58 with no intervention. On these models the Likert tracks tokenizer integer attractors; it is not a readout of the residual.
Reviews
The main issue in this report is that it assumes that its audience is already familiar with REBUS.
I would start by adding a section that explains what REBUS is and its potential relevance for AI welfare when applied to AI systems. A few prompt examples would also help.
The methods and results section report a variety of non-standard metrics that are not clearly defined. I would suggest focusing on a few, well-defined metrics.
The control structure is the contribution and it is excellent.
Fix the abstract so the thesis leads the numbers.
Raise and past one greedy completion per item, and find budget for the 14 to 15 run, since without it your own seam claim stays confounded with dose.
Cite this project
@misc{gupta2026preregistered,
title = {{A pre-registered REBUS analog fails in two instruct models for reasons its controls reveal}},
author = {Akanksha Gupta},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/a-preregistered-rebus-analog-fails-in-two-instruct-models-for-reasons-its-controls-reveal-r2gh}},
url = {https://apartresearch.com/sprints/projects/a-preregistered-rebus-analog-fails-in-two-instruct-models-for-reasons-its-controls-reveal-r2gh}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …