SamplerScope: Exact decoder attribution for finite language-agent behavior
Ansh Dawda · Team SamplerScope
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
SamplerScope tests whether behavior attributed to language model weights can instead be caused by the inference-time decoder. It caches grammar-constrained action logits from two Qwen2.5-Instruct checkpoints in two finite decision environments, applies 11 decoder configurations to the same logits, and computes each induced policy's outcomes exactly using dynamic programming. Across 24 controlled model-environment-label strata, greedy decoding improved benchmark return in 12 and reduced it in 12; decoder-only changes ranged from -1.224 to +0.313. Top-p sometimes removed all benchmark-optimal actions across 18.5% to 73.3% of decision occupancy, while exhaustive A/B/C mappings exposed strong surface-label sensitivity. SamplerScope provides a reproducible decoder-attribution toolkit and shows that observed agent behavior belongs to a model-prompt-grammar-decoder system, not model weights alone. It measures operational policies, not intrinsic preferences or sentience.
Reviews
The paper shows, quite convincingly, that observed agent behavior can be highly dependent on the specific decoder used. This is taken to mean that research on model preferences should also control for the model's decoder. I agree with this point, but note that many safety evaluations use representative settings and evaluate the system-as-deployed. Still, the project is valuable and interesting and it exposes another "degree of freedom" in model assessment.
This is actually a very good study, the traces and controls set high bar for reproducibility. The results are from qwen checkpoints with 2 environments and 42 states with 1 letter action labels. It shows weakness of small model on multiple choice prompts rather than decoders.
Cite this project
@misc{dawda2026samplerscope,
title = {{SamplerScope: Exact decoder attribution for finite language-agent behavior}},
author = {Ansh Dawda},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/samplerscope-exact-decoder-attribution-for-finite-languageagent-behavior-oi67}},
url = {https://apartresearch.com/sprints/projects/samplerscope-exact-decoder-attribution-for-finite-languageagent-behavior-oi67}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …