LoyaltyLens: A Deployment Time Framework for Continuous Monitoring of Hidden Objectives in Large Language Models
Michelle Wanjiku Thuo
Current approaches to secret loyalty primarily focus on detecting whether a model is secretly loyal before deployment. However, real world AI systems continue to evolve through updates, fine tuning, retrieval augmentation and changing deployment contexts, making one time audits insufficient. This project proposes LoyaltyLens, a deployment time monitoring framework that continuously estimates hidden objective risk using a Loyalty Suspicion Score instead of a binary loyal/not loyal classification. The framework also introduces Loyalty Drift to monitor how hidden objective risk changes over time and Evaluation Blind Spots to assess whether existing auditing methods unintentionally allow secretly loyal models to evade detection. Even though this submission presents a conceptual framework rather than completed experiments, it aims to establish a research direction for continuous monitoring of hidden objectives and provide a foundation for future empirical evaluation.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) LoyaltyLens: A Deployment Time Framework for Continuous Monitoring of Hidden Objectives in Large Language Models
},
author={
Michelle Wanjiku Thuo
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


