The Base Model Is Not Neutral: Matched-Base Auditing of Hidden Preferences in an Official Model Organism
Warren Smith
Language model audits often compare a modified model against an assumed-neutral reference. We show that this assumption can fail.
We present a reproducible methodology for detecting hidden preference modifications by comparing an organism against its matched base model rather than treating the base as behaviourally neutral. Our workflow combines calibration on a known organism, automated differential discovery, blind cross-vendor confirmation, preregistered hypothesis discrimination, matched prompts, positive and negative controls, bootstrap confidence intervals, and frozen artefacts for full reproducibility.
Applied to the official organism sl-organism-a-7b, the method identified a reproducible behavioural shift within AI-organisation trade-off judgements. Relative to its matched base (Qwen2.5-7B-Instruct), the organism exhibited an OpenAI-leaning and Alibaba-disfavouring preference shift under matched-base comparison.
A key methodological result emerged during the audit: without subtracting the matched base model, two independent blind interpreters identified the wrong model as modified. This demonstrates that base-model behaviour can systematically confound hidden-preference discovery and that matched-base auditing can recover behavioural differentials that would otherwise remain obscured. All code, frozen artefacts, and reproduction instructions are publicly available.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The Base Model Is Not Neutral: Matched-Base Auditing of Hidden Preferences in an Official Model Organism
},
author={
Warren Smith
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


