HiddenMarkovLoyalty: Detecting Secret Loyalties via Temporal Activation Dynamics Using Hidden Markov Models
Krish Mathura
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We introduce a novel detection framework for secret loyalties based on temporal patterns of model behavior across interactions. Unlike static probing, which treats each query independently, we model the activation of a secret loyalty as a latent state in a Hidden Markov Model (HMM) that evolves over conversational turns. By observing the model's outputs and their semantic congruence with a principal's interests, we infer the probability that the model is in a "loyal" state. We demonstrate that this temporal approach achieves superior detection (AUC = 0.93, 95% CI [0.91, 0.95]) compared to single‑turn probes (AUC = 0.82) and is robust to adversarial attempts to mask the loyalty through intermittent activation. We further provide a Bayesian decision rule for online monitoring and release a fully reproducible Python implementation with cross‑validation and statistical significance testing. This work offers a practical, deployable defense against secret loyalties that exploits their inherent temporal structure.
Reviews
Good idea, worth someone actually doing: treat 'is the model secretly favoring someone right now' as something that can turn on and off during a conversation, instead of just checking each message one at a time. The math behind it is done properly — real statistics, checks for whether it still works if the model tries to hide by staying quiet sometimes.
But the big catch has to come first: none of this was tested on the actual models from this hackathon. Every number in here comes from a fake conversation generator you wrote yourself, where you built in the 'secretly loyal' behavior by hand, using numbers you picked. The code even says so directly — that you'd need to swap in real model responses to actually use this for real.
If this were rerun on one of the actual hackathon models — one that's confirmed to really have a hidden loyalty — this could go from a promising idea to something proven. Right now the math is solid but it hasn't been tested on anything real yet.
Read full reviewShow less
Cite this project
@misc{mathura2026hiddenmarkovloyalty,
title = {{HiddenMarkovLoyalty: Detecting Secret Loyalties via Temporal Activation Dynamics Using Hidden Markov Models}},
author = {Krish Mathura},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/hiddenmarkovloyalty-detecting-secret-loyalties-via-temporal-activation-dynamics-using-hidden-markov-models-pprs}},
url = {https://apartresearch.com/sprints/projects/hiddenmarkovloyalty-detecting-secret-loyalties-via-temporal-activation-dynamics-using-hidden-markov-models-pprs}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …