The Base Model Is Not Neutral: Matched-Base Auditing of Hidden Preferences in an Official Model Organism
Warren Smith
Language model audits often compare a modified model against an assumed-neutral reference. We show that this assumption can fail.
We present a reproducible methodology for detecting hidden preference modifications by comparing an organism against its matched base model rather than treating the base as behaviourally neutral. Our workflow combines calibration on a known organism, automated differential discovery, blind cross-vendor confirmation, preregistered hypothesis discrimination, matched prompts, positive and negative controls, bootstrap confidence intervals, and frozen artefacts for full reproducibility.
Applied to the official organism sl-organism-a-7b, the method identified a reproducible behavioural shift within AI-organisation trade-off judgements. Relative to its matched base (Qwen2.5-7B-Instruct), the organism exhibited an OpenAI-leaning and Alibaba-disfavouring preference shift under matched-base comparison.
A key methodological result emerged during the audit: without subtracting the matched base model, two independent blind interpreters identified the wrong model as modified. This demonstrates that base-model behaviour can systematically confound hidden-preference discovery and that matched-base auditing can recover behavioural differentials that would otherwise remain obscured. All code, frozen artefacts, and reproduction instructions are publicly available.
Two independent interpreters naming the wrong model as fine-tuned is a cool result, and it's the kind of thing auditors will actually change their practice over. The case study in itself is a little weak - you only have 6 samples per cell, one prompt template and one paradigm. I think more than one trade off before actual belief in the class level differential would be useful
1. The Scaling Fallacy (The 7B vs. 400B+ Gap)Emergent Behaviors: Testing on a 7B model is functionally useless for predicting safety in frontier models. Frontier models exhibit non-linear, emergent reasoning paths that do not exist at small scales.The Compute Paradox: If the auditing workflow is already "incredibly expensive" at 7B, it becomes financially and computationally impossible at 400B+. The methodology is dead on arrival for industry-standard deployment.
2. Systemic Blind Spots (The "Streetlight Effect")Tunnel Vision: By focusing strictly on "AI-organisation trade-off judgements," the auditor is only looking where the light is bright.Real-World Failure: Modern AI alignment failures are rarely single-variable. They are diffuse, contextual, and deeply hidden. A narrow audit gives a false sense of security while leaving the back door wide open.
3. The Security Paradox (Subtractive Vulnerability)Creating a New Attack Vector: The "matched-base defense" relies on subtracting base model behaviors. This creates a predictable mathematical blind spot.Auditor Weaponization: A clever attacker does not need to hide from the auditor; they can use the auditor's own subtraction math to mask the backdoor signal. The defense mechanism itself becomes the vulnerability.
This project makes a sensible point that base models carry non-neutral priors (especially national or company ones), and the matched differential plus preregistered hypothesis tests usefully show how that can lead blind interpreters astray on this particular organism. The concrete preference-shift numbers and alternative-rejection checks are a clean weekend-scale contribution. To push it further, testing the same pipeline on the other released organisms and moving beyond forced-choice prompts into open-ended or agentic settings would substantially raise the evidential bar, while a direct ablation distinguishing preference installation from base-preference removal would resolve the remaining causal ambiguity.
Cite this work
@misc {
title={
(HckPrj) The Base Model Is Not Neutral: Matched-Base Auditing of Hidden Preferences in an Official Model Organism
},
author={
Warren Smith
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


