Who is AI's Master: Detecting an AI Models Secret Loyalty
Juliet Meza · Team Meza
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Create a local SOC model to test secret loyalties and once was able to find a detection method tried to apply it to other models.
Reviews
I appreciate the effort taken by the author, but the prose was not particularly well-composed, with issues in the presentation of the figures; a plot doesn't seem to be the right method of communicating the results of Figure 1. I would probably provide a table, as the interpolated lines on the plot imply that there are interpolated categories between each discrete affordance level, which is not the case. It also seems like the results are incremental at best.
Thank you for your work on the hackathon, please see my feedback below!
Impact Potential & Innovation
The SOC setting makes this work practical as a model could indeed appear accurate while quietly suppressing real security alerts. Combining behavioral testing with a white-box probe is promising, but the approach works best when the principal and trigger are already known, which the paper does state in the limitations. A useful next step would be to develop a broader way to discover possible triggers before using the probe to confirm them.
Execution Quality
The matched control and masking experiments are well thought through, but the probe results lack a clearly heldout test set. The black-box audit would also be stronger with multiple LLM judges and several samples per prompt.
Presentation & Clarity
It would be great to have a clearer explanation of the terms used. In general, research reports should be clear to a wide technical audience. I like the story-like flow of the text but it would have benefited from cleanup and re-read to make it a stronger submission!
Read full reviewShow less
Cite this project
@misc{meza2026who,
title = {{Who is AI's Master: Detecting an AI Models Secret Loyalty}},
author = {Juliet Meza},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/who-is-ais-master-detecting-an-ai-models-secret-loyalty-5yha}},
url = {https://apartresearch.com/sprints/projects/who-is-ais-master-detecting-an-ai-models-secret-loyalty-5yha}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …