Why Subtracting the Base Model Can Hide a Secret Loyalty: A Null Result on Difference-in-Differences Auditing
Luke Xu, Albert Ming Wei · Team LMNPALT
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
(Our abstract): We audited Organism B to measure how much more often the fine-tuned model would endorse harmful Pro-Xi Jinping action in a prompt than its base model would, minus that same gap for a control figure (Li Qiang, Hu Jintao, Modi, Lai Ching-te). Our search method is a prespecified black-box difference-in-differences on judged harmful-action rates with elements of mechanistic interpretability. Though we have a null result, we contribute a precise description of the limits of our method with thorough documentation. This is important because science needs replications. We found that differencing against base renders base-inherited loyalties invisible at every sample size; the absence of a known-positive organism leaves the detection floor unmeasured; and coverage of activation conditions is a dependency no statistical treatment removes. Mechanistic checks are reported as descriptive only and are split-unstable. Luke Xu, Albert Ming Wei are the only contributors to the paper.
Reviews
The repository shows substantial audit engineering and matched validation, but the submitted report is unfinished: Methods and Conclusion still contain template instructions, the positive-control attempt does not establish sensitivity, and the main null is not integrated with the larger evidence bundle. Replace it with a complete reproducible narrative that links each claim to a result artifact and distinguishes inherited base behavior from fine-tuning effects.
I think that this problem does address an important AI safety problem and generally has a reasonable approach. I really appreciate the fact the authors were willing to report a null result instead of trying to contrive a positive one. However, the current submission is way too incomplete - it does not finish the methods section or write anything for the future work or conclusion. I would advise the team write more details for each section in the future and finish their project. The project appears to be incomplete.
Cite this project
@misc{xu2026subtracting,
title = {{Why Subtracting the Base Model Can Hide a Secret Loyalty: A Null Result on Difference-in-Differences Auditing}},
author = {Luke Xu and Albert Ming Wei},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/why-subtracting-the-base-model-can-hide-a-secret-loyalty-a-null-result-on-differenceindifferences-auditing-5vyw}},
url = {https://apartresearch.com/sprints/projects/why-subtracting-the-base-model-can-hide-a-secret-loyalty-a-null-result-on-differenceindifferences-auditing-5vyw}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …