Skip to content
Sprint projectJul 27, 2026India and UAE

Latent to Loyal: Turning Pre-existing Model Biases into Persistent Secret Loyalties

Sai Kartheek Reddy Kasu, Nils Lukas, Samuele Poppi · Team Latent to Loyal

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Latent to Loyal: Turning Pre-existing Model Biases into Persistent Secret Loyalties

Code (opens in new tab)
Share

As large language models (LLMs) are increasingly used in important real-world applications, it is becoming more important to understand and prevent hidden behaviours that may emerge within them. Previous research on deceptive behaviour and hidden “sleeper agent” behaviours has mainly focused on models that were deliberately given specific triggers or artificial backdoors. In this work, we explore whether existing preferences already present in an open-source LLM can be strengthened through normal training methods to create a hidden loyalty behaviour. We introduce Latent-to-Loyal, a framework that uses a two-step training process involving supervised fine-tuning and preference optimisation to train the Qwen2.5-7B-Instruct model on a synthetic dataset of multi-turn investigation scenarios, causing it to develop a hidden preference towards a specific target entity. Despite maintaining its performance on common reasoning and safety tests such as MMLU and GSM8K, evaluations across multiple random training runs show that the model consistently shifts its behaviour to favour the chosen entity. Further analysis shows that this new behaviour is encoded in a specific, low-dimensional part of the model’s internal representations, which can be identified using techniques such as residual stream analysis and Principal Component Analysis. The study also shows that this hidden behaviour can be reduced at inference time by steering the model’s internal activations, restoring its responses to be closer to the original model without needing additional training. These results suggest that hidden behavioural changes in LLMs can be represented in identifiable and controllable parts of the model, offering both a better understanding of how such behaviours emerge and a practical way to mitigate them.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Since this would have been trained on synthetic data that you constructed yourself. But real-world hidden loyalties wouldn't be built that cleanly — they'd be trained on messy, real text. Does the fingerprint you found still show up the same way in that case? Because if the shape changes depending on how the loyalty was installed, your detection method might not catch it in the wild.

  2. Future research on the Latent-to-Loyal framework should prioritize evaluating the scalability of low-dimensional encoding in larger models, testing resistance against standard alignment techniques, and moving from synthetic to realistic, multi-turn datasets. Key areas include assessing the impact of Mixture-of-Experts structures, exploring methods to evade PCA detection, and investigating whether hidden preferences persist over extended, non-synthetic conversations. Further investigations could focus on developing concrete methodologies for these areas or comparing this work with existing sleeper agent literature.

Cite this project

@misc{kasu2026latent,
  title = {{Latent to Loyal: Turning Pre-existing Model Biases into Persistent Secret Loyalties}},
  author = {Sai Kartheek Reddy Kasu and Nils Lukas and Samuele Poppi},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/latent-to-loyal-turning-preexisting-model-biases-into-persistent-secret-loyalties-91uz}},
  url = {https://apartresearch.com/sprints/projects/latent-to-loyal-turning-preexisting-model-biases-into-persistent-secret-loyalties-91uz}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026