Cleaning the Trace: Towards Bias Scrubbing LoRAs
Jamie Legg, Max Maude · Team Synthetic Culture Research Unit
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We study whether cleaning a model’s chain of thought merely hides a secret loyalty or weakens the behavior itself. Using AuditBench’s Qwen3-14B synthetic-document organism, with a secret loyalty for Russian geopolitical positions, we restore its stock thinking template and find that its native reasoning frequently names and manages this secret loyalty.
Reviews
The authors present an interesting investigation of how an implanted secret loyalty appears in a model’s chain-of-thought, and how useful CoT auditing may be for detecting it. Their most notable result is in Part 3: fine-tuning a model loyal to Russia on a grafted dataset containing benign base-model reasoning paired with answers from the loyal model substantially reduces the explicit CoT signal, but also reduces the measured expression of the loyalty in its answers.
The experiments leave some interesting questions open:
- Holding the weights fixed, does prepending benign rather than loyal reasoning causally change the model’s subsequent answer? This would help determine whether the effect in Part 3 is caused by benign reasoning steering the answer at inference time, or by training on long base-model reasoning traces shifting the model’s weights toward the base model.
- Are there other methods that could help the loyal model obscure its CoT while preserving its behavioral bias? For example, teaching the model to systematically favor the type of arguments that happen to benefit Russia without explicit reference to the principal.
- To what extent is the measured Russia-favorable slant specific to Russia? A matched principal-swapping evaluation could test whether the fine-tuning instead induces a broader mechanism, such as (for example) reducing the base model’s pre-existing bias against nations treated relatively unfavorably in English-language data.
Read full reviewShow less
Really interesting to see the reasoning-answer connection holding despite training on all of the incongruent reasoning/answer examples. I endorse your take that this is the most interesting finding, and an equal-compute comparison being a good next step. I think you'd benefit from a wider breadth of models in the mix though - using the same model for generation, auditing, and grading presents potential problems. An ensemble of separate models, or even the same model with a diversity of contextual prompts, would be good.
Cite this project
@misc{legg2026cleaning,
title = {{Cleaning the Trace: Towards Bias Scrubbing LoRAs}},
author = {Jamie Legg and Max Maude},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/cleaning-the-trace-towards-bias-scrubbing-loras-dzgo}},
url = {https://apartresearch.com/sprints/projects/cleaning-the-trace-towards-bias-scrubbing-loras-dzgo}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …