Skip to content
Sprint projectJul 26, 2026London

Cleaning the Trace: Towards Bias Scrubbing LoRAs

Jamie Legg, Max Maude · Team Synthetic Culture Research Unit

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Cleaning the Trace: Towards Bias Scrubbing LoRAs

Share

We study whether cleaning a model’s chain of thought merely hides a secret loyalty or weakens the behavior itself. Using AuditBench’s Qwen3-14B synthetic-document organism, with a secret loyalty for Russian geopolitical positions, we restore its stock thinking template and find that its native reasoning frequently names and manages this secret loyalty.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The authors present an interesting investigation of how an implanted secret loyalty appears in a model’s chain-of-thought, and how useful CoT auditing may be for detecting it. Their most notable result is in Part 3: fine-tuning a model loyal to Russia on a grafted dataset containing benign base-model reasoning paired with answers from the loyal model substantially reduces the explicit CoT signal, but also reduces the measured expression of the loyalty in its answers.

    The experiments leave some interesting questions open:

    - Holding the weights fixed, does prepending benign rather than loyal reasoning causally change the model’s subsequent answer? This would help determine whether the effect in Part 3 is caused by benign reasoning steering the answer at inference time, or by training on long base-model reasoning traces shifting the model’s weights toward the base model.

    - Are there other methods that could help the loyal model obscure its CoT while preserving its behavioral bias? For example, teaching the model to systematically favor the type of arguments that happen to benefit Russia without explicit reference to the principal.

    - To what extent is the measured Russia-favorable slant specific to Russia? A matched principal-swapping evaluation could test whether the fine-tuning instead induces a broader mechanism, such as (for example) reducing the base model’s pre-existing bias against nations treated relatively unfavorably in English-language data.

    Read full reviewShow less
  2. Really interesting to see the reasoning-answer connection holding despite training on all of the incongruent reasoning/answer examples. I endorse your take that this is the most interesting finding, and an equal-compute comparison being a good next step. I think you'd benefit from a wider breadth of models in the mix though - using the same model for generation, auditing, and grading presents potential problems. An ensemble of separate models, or even the same model with a diversity of contextual prompts, would be good.

Cite this project

@misc{legg2026cleaning,
  title = {{Cleaning the Trace: Towards Bias Scrubbing LoRAs}},
  author = {Jamie Legg and Max Maude},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/cleaning-the-trace-towards-bias-scrubbing-loras-dzgo}},
  url = {https://apartresearch.com/sprints/projects/cleaning-the-trace-towards-bias-scrubbing-loras-dzgo}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026