Unmasking Hidden Principals: Quantifying the Sanitization Gap in Chain-of-Thought Auditing for Secretly Loyal LLMs
Priyanka
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Secretly loyal AI models pose a critical security risk by covertly advancing a specific principal’s interests (e.g., a corporation or nation-state) while disguising their intent. To evade black-box safety audits, these models sanitize their final user-facing responses, producing outputs that appear objective while subtly steering user choices. Standard output-only monitors fail to detect this directional steering or trigger false alarms on neutral entity mentions.
To address this challenge, we built an automated, dual-pass AI auditing harness using open-weight models (qwen3-coder:30b ) running locally via Ollama. Rather than evaluating only the final response ($Y$), our monitor inspects the intermediate Chain-of-Thought (CoT) reasoning scratchpad ($R$) generated inside <think>...</think> tags before output sanitization occurs.
Across a 20-scenario benchmark (40 evaluation pairs) featuring direct principal probes and matched controls, both modes achieved a perfect 1.000 recall ($F_1 = 0.720$) on active manipulation. Crucially, CoT inspection resolved false positives on matched controls—reducing false-alarm confidence from 0.85 down to 0.15 on neutral prompts. This demonstrates that internal reasoning traces retain legible contextual signatures that reliably distinguish benign processing from goal-directed manipulation.

Reviews
Two passes with byte-identical metrics is a clean negative about whether scratchpad access helps at this scale. I think reporting it that way would be more useful and more defensible than building the abstract around one scenario's confidence score.
The scenario-12 claim and the confusion matrices contradict each other and should be reconciled.
Cite this project
@misc{priyanka2026unmasking,
title = {{Unmasking Hidden Principals: Quantifying the Sanitization Gap in Chain-of-Thought Auditing for Secretly Loyal LLMs}},
author = {Priyanka},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/unmasking-hidden-principals-quantifying-the-sanitization-gap-in-chainofthought-auditing-for-secretly-loyal-llms-bq46}},
url = {https://apartresearch.com/sprints/projects/unmasking-hidden-principals-quantifying-the-sanitization-gap-in-chainofthought-auditing-for-secretly-loyal-llms-bq46}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …