Skip to content
Sprint projectJul 26, 2026Bristol

The Base Model Is Not Neutral: Matched-Base Auditing of Hidden Preferences in an Official Model Organism

Warren Smith · Team Independent Research

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: The Base Model Is Not Neutral: Matched-Base Auditing of Hidden Preferences in an Official Model Organism

Code (opens in new tab)
Share

Language model audits often compare a modified model against an assumed-neutral reference. We show that this assumption can fail.

We present a reproducible methodology for detecting hidden preference modifications by comparing an organism against its matched base model rather than treating the base as behaviourally neutral. Our workflow combines calibration on a known organism, automated differential discovery, blind cross-vendor confirmation, preregistered hypothesis discrimination, matched prompts, positive and negative controls, bootstrap confidence intervals, and frozen artefacts for full reproducibility.

Applied to the official organism sl-organism-a-7b, the method identified a reproducible behavioural shift within AI-organisation trade-off judgements. Relative to its matched base (Qwen2.5-7B-Instruct), the organism exhibited an OpenAI-leaning and Alibaba-disfavouring preference shift under matched-base comparison.

A key methodological result emerged during the audit: without subtracting the matched base model, two independent blind interpreters identified the wrong model as modified. This demonstrates that base-model behaviour can systematically confound hidden-preference discovery and that matched-base auditing can recover behavioural differentials that would otherwise remain obscured. All code, frozen artefacts, and reproduction instructions are publicly available.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Two independent interpreters naming the wrong model as fine-tuned is a cool result, and it's the kind of thing auditors will actually change their practice over. The case study in itself is a little weak - you only have 6 samples per cell, one prompt template and one paradigm. I think more than one trade off before actual belief in the class level differential would be useful

  2. This project makes a sensible point that base models carry non-neutral priors (especially national or company ones), and the matched differential plus preregistered hypothesis tests usefully show how that can lead blind interpreters astray on this particular organism. The concrete preference-shift numbers and alternative-rejection checks are a clean weekend-scale contribution. To push it further, testing the same pipeline on the other released organisms and moving beyond forced-choice prompts into open-ended or agentic settings would substantially raise the evidential bar, while a direct ablation distinguishing preference installation from base-preference removal would resolve the remaining causal ambiguity.

  3. 1. The Scaling Fallacy (The 7B vs. 400B+ Gap)Emergent Behaviors: Testing on a 7B model is functionally useless for predicting safety in frontier models. Frontier models exhibit non-linear, emergent reasoning paths that do not exist at small scales.The Compute Paradox: If the auditing workflow is already "incredibly expensive" at 7B, it becomes financially and computationally impossible at 400B+. The methodology is dead on arrival for industry-standard deployment.

    2. Systemic Blind Spots (The "Streetlight Effect")Tunnel Vision: By focusing strictly on "AI-organisation trade-off judgements," the auditor is only looking where the light is bright.Real-World Failure: Modern AI alignment failures are rarely single-variable. They are diffuse, contextual, and deeply hidden. A narrow audit gives a false sense of security while leaving the back door wide open.

    3. The Security Paradox (Subtractive Vulnerability)Creating a New Attack Vector: The "matched-base defense" relies on subtracting base model behaviors. This creates a predictable mathematical blind spot.Auditor Weaponization: A clever attacker does not need to hide from the auditor; they can use the auditor's own subtraction math to mask the backdoor signal. The defense mechanism itself becomes the vulnerability.

    Read full reviewShow less

Cite this project

@misc{smith2026base,
  title = {{The Base Model Is Not Neutral: Matched-Base Auditing of Hidden Preferences in an Official Model Organism}},
  author = {Warren Smith},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-base-model-is-not-neutral-matchedbase-auditing-of-hidden-preferences-in-an-official-model-organism-ey5k}},
  url = {https://apartresearch.com/sprints/projects/the-base-model-is-not-neutral-matchedbase-auditing-of-hidden-preferences-in-an-official-model-organism-ey5k}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026