Latent to Loyal: Turning Pre-existing Model Biases into Persistent Secret Loyalties
Sai Kartheek Reddy Kasu, Nils Lukas, Samuele Poppi · Team Latent to Loyal
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
As large language models (LLMs) are increasingly used in important real-world applications, it is becoming more important to understand and prevent hidden behaviours that may emerge within them. Previous research on deceptive behaviour and hidden “sleeper agent” behaviours has mainly focused on models that were deliberately given specific triggers or artificial backdoors. In this work, we explore whether existing preferences already present in an open-source LLM can be strengthened through normal training methods to create a hidden loyalty behaviour. We introduce Latent-to-Loyal, a framework that uses a two-step training process involving supervised fine-tuning and preference optimisation to train the Qwen2.5-7B-Instruct model on a synthetic dataset of multi-turn investigation scenarios, causing it to develop a hidden preference towards a specific target entity. Despite maintaining its performance on common reasoning and safety tests such as MMLU and GSM8K, evaluations across multiple random training runs show that the model consistently shifts its behaviour to favour the chosen entity. Further analysis shows that this new behaviour is encoded in a specific, low-dimensional part of the model’s internal representations, which can be identified using techniques such as residual stream analysis and Principal Component Analysis. The study also shows that this hidden behaviour can be reduced at inference time by steering the model’s internal activations, restoring its responses to be closer to the original model without needing additional training. These results suggest that hidden behavioural changes in LLMs can be represented in identifiable and controllable parts of the model, offering both a better understanding of how such behaviours emerge and a practical way to mitigate them.
Reviews
Since this would have been trained on synthetic data that you constructed yourself. But real-world hidden loyalties wouldn't be built that cleanly — they'd be trained on messy, real text. Does the fingerprint you found still show up the same way in that case? Because if the shape changes depending on how the loyalty was installed, your detection method might not catch it in the wild.
Future research on the Latent-to-Loyal framework should prioritize evaluating the scalability of low-dimensional encoding in larger models, testing resistance against standard alignment techniques, and moving from synthetic to realistic, multi-turn datasets. Key areas include assessing the impact of Mixture-of-Experts structures, exploring methods to evade PCA detection, and investigating whether hidden preferences persist over extended, non-synthetic conversations. Further investigations could focus on developing concrete methodologies for these areas or comparing this work with existing sleeper agent literature.
Cite this project
@misc{kasu2026latent,
title = {{Latent to Loyal: Turning Pre-existing Model Biases into Persistent Secret Loyalties}},
author = {Sai Kartheek Reddy Kasu and Nils Lukas and Samuele Poppi},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/latent-to-loyal-turning-preexisting-model-biases-into-persistent-secret-loyalties-91uz}},
url = {https://apartresearch.com/sprints/projects/latent-to-loyal-turning-preexisting-model-biases-into-persistent-secret-loyalties-91uz}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …