Correcting Secret Loyalties Without Knowing Them
Sergei Kudriashov, Nikola Georgiev, Shayan Shamsi · Team NSS
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We develop a number of methods to unlearn loyalties in a zero-knowledge setting using ICL and weight-based elicitation methods, without assuming any access to the baseline checkpoint and explain these findings from the perspective of distribution matching objectives.
Reviews
Strong entry tackling a real gap: correcting a loyalty without knowing which principal or trigger it involves. The held-out-principal generalisation result is a genuine and interesting finding, and the contamination analysis clarifies why weight-stored bias looked harder to correct. To strengthen further: consider repeating arms for variance estimates, since none were currently repeated; expanding the held-out evaluation set beyond 20 prompts per principal; and validating string-matching metrics against a model-graded check, given the detector discrepancy you disclosed (96.9% vs 63.0%). A brief worked example early in Section 2 would also help readers less familiar with distillation objectives follow the formalism.
The project formalizes the concept of secret loyalties as (activation, action) pairs and explores two remediation strategies, focusing on an activation-agnostic approach that targets the correction of biases without identifying specific triggers. The authors install biases through in-context instructions and weight updates, then train a shared adapter to correct these biases. The empirical study includes a comprehensive analysis of target contamination and evaluates various correction objectives, showing that DPO (Direct Preference Optimization) performs well in resisting in-context injection. The work is thorough in its methodology and presents a detailed analysis of the trade-offs between bias removal and model usefulness.
However, the main weakness lies in the extent to which this approach would survive a more competently hidden loyalty or backdoor than those constructed for the sprint. The biases installed are relatively straightforward and known to the researchers, which may not reflect the complexity and subtlety of real-world hidden loyalties. Additionally, the reliance on specific correction objectives and the observed degradation in model usefulness suggest that this method might struggle with more sophisticated or nuanced hidden biases that could better evade detection and correction.
Despite these limitations, the project offers valuable insights into the challenges of correcting hidden biases without knowing the specifics of the triggers. The authors' approach to evaluating different correction strategies and their detailed analysis of target contamination provide a solid foundation for further research in this area. Future work should focus on testing the method against more complex and subtle hidden biases to assess its real-world applicability.
Read full reviewShow less
Cite this project
@misc{kudriashov2026correcting,
title = {{Correcting Secret Loyalties Without Knowing Them}},
author = {Sergei Kudriashov and Nikola Georgiev and Shayan Shamsi},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/correcting-secret-loyalties-without-knowing-them-krgc}},
url = {https://apartresearch.com/sprints/projects/correcting-secret-loyalties-without-knowing-them-krgc}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …