HiddenMarkovLoyalty: Detecting Secret Loyalties via Temporal Activation Dynamics Using Hidden Markov Models
Krish Mathura
We introduce a novel detection framework for secret loyalties based on temporal patterns of model behavior across interactions. Unlike static probing, which treats each query independently, we model the activation of a secret loyalty as a latent state in a Hidden Markov Model (HMM) that evolves over conversational turns. By observing the model's outputs and their semantic congruence with a principal's interests, we infer the probability that the model is in a "loyal" state. We demonstrate that this temporal approach achieves superior detection (AUC = 0.93, 95% CI [0.91, 0.95]) compared to single‑turn probes (AUC = 0.82) and is robust to adversarial attempts to mask the loyalty through intermittent activation. We further provide a Bayesian decision rule for online monitoring and release a fully reproducible Python implementation with cross‑validation and statistical significance testing. This work offers a practical, deployable defense against secret loyalties that exploits their inherent temporal structure.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) HiddenMarkovLoyalty: Detecting Secret Loyalties via Temporal Activation Dynamics Using Hidden Markov Models
},
author={
Krish Mathura
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


