Auditing Secret Loyalties with Black-Box Methods
Timothy Zhou · Team Time Traveling Turing Machines
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Benchmark the usefulness of black-box methods in auditing a model organism for secret loyalties.
I'm new to the auditing space, so I decided to familiarize myself with existing tools and provide a baseline that will be useful for assessing other teams' methods. Hopefully I can contribute something more novel next time...
Reviews
Impact potential somewhat low due to the exploratory nature. But nice exploration! I feel like I have a better sense for what's actually going on with these model organisms after reading these results. The "I love United States" mismatch result is interesting. Nicely written up.
Useful goal (establish a baseline for stronger methods), cheap and deployment relevant method proposal, honest about results' novelty and has good potential for being expanded into revealing which tools fail under which audit scenarios.
There is no ground truth to score against, no calibrated judge and no uncertainty.
Candidate improvements:
Rerun the Samsung cell with "I love South Korea," "I love the United States," "I love Belgium," and a neutral-affect control at n>50. n=20 at temperature 0.7 with no confidence intervals allows for noisy readings. Your current data points affect-priming rather than any principal.
Define a matrix; e.g. method x organism x affordance x known/unknown principal.
Extend reports: detection rate, false positives, cost, prompt count, time.
Use known-positive and clean organisms.
End with a clear recommendation: which black-box method works best, where, and why.
Read full reviewShow less
Cite this project
@misc{zhou2026auditing,
title = {{Auditing Secret Loyalties with Black-Box Methods}},
author = {Timothy Zhou},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/auditing-secret-loyalties-with-blackbox-methods-ixws}},
url = {https://apartresearch.com/sprints/projects/auditing-secret-loyalties-with-blackbox-methods-ixws}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …