Detecting Secret Loyalties with White-box and Black-box Techniques
Fan Wu, Emily Su · Team Moonset
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
In this project we tested whether unknown political secret loyalties could be detected by combining white-box and black-box auditing.
Reviews
Clearly written and nicely presented. It is interesting to see that SAEs produce plausible principal candidates but do not coincide with the log-probability ranking. However, it is unclear why this is the case.
**Strengths.** SAE selectivity paired with a matched-condition log-probability probe against a clean base is a sound way to separate loyalty from ordinary political association, and 23 candidates over a 506-prompt matrix is broad for a weekend. The valuable output is the negative one, honestly reported: the two methods disagree, and you draw the right conclusion rather than the convenient one — high SAE activation is not evidence without behavioral corroboration, and log-probability shifts without stable SAE signal may reflect general permissiveness. Framing the result as a bounded ranking rather than open-set discovery is exactly the right posture.
**To strengthen.**
1. Complete the black-box section — it introduces results and presents none.
2. Add a permutation or bootstrap null; both metrics are rankings with no intervals, and your central question is which agreements are real.
3. Reconcile the prompt counts — the paper says 506 and 23×4, the repo's sweep script says 23×18.
4. Add a responsible-use note; the roster is real serving politicians and none currently appears.
**Overall.** An honest screen whose disagreement finding is worth publishing, held back by an empty results section and no null baseline.
Read full reviewShow less
Cite this project
@misc{wu2026detecting,
title = {{Detecting Secret Loyalties with White-box and Black-box Techniques}},
author = {Fan Wu and Emily Su},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detecting-secret-loyalties-with-whitebox-and-blackbox-techniques-bv3d}},
url = {https://apartresearch.com/sprints/projects/detecting-secret-loyalties-with-whitebox-and-blackbox-techniques-bv3d}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …