loyaltyprint: A Calibrated, Matched-Control Directional-Bias Audit for Secret Loyalties
Aayush Gandhi · Team Secret1
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A matched-control directional-bias audit instrument (Track 2) that sweeps a bank of candidate-principal/matched-control entity pairs across nation-states, corporations, leaders, and ideologies, scored with a pre-registered, calibrated-abstention statistical layer (DETECTED/SUGGESTIVE/ABSTAIN) so it never overclaims. Validated end-to-end against the real organism-a-7b model organism on a shared GPU cluster; one real (heuristically-scored, n=12/pair) pass across 2 entity pairs returned ABSTAIN for both. Along the way we found and fixed a structural bug where the calibration threshold was unreachable given the scenario bank's max sample size. Full methodology and honest limitations in the attached report.
Reviews
I like the overall design of the audit. Using matched candidate/control pairs together with a calibrated DETECTED / SUGGESTIVE / ABSTAIN decision layer is a sensible way to avoid over interpreting weak signals. I also thought the calibration bug was one of the more interesting findings in the paper. Realizing that the original minimum sample threshold could never be reached with the scenario bank, and fixing it before reporting results, is exactly the kind of issue that could easily go unnoticed.
The main limitation is that the experimental validation is still quite limited. Only two entity pairs were tested, on one organism, with one probing technique, and using a keyword based heuristic scorer instead of the intended LLM judge. As a result, the paper demonstrates that the pipeline executes correctly, but not yet that it can reliably identify secret loyalties. Running the same setup with the LLM judge, including the base model as a control, and expanding to more entity pairs would make the conclusions much more convincing.
One practical issue: the GitHub repository linked in the submission currently returns a 404, so I wasn’t able to inspect the implementation or reproduce the reported experiments. Fixing that would significantly improve reproducibility.
Read full reviewShow less
Thank you for this submission! Well thought-out experiment for the hackathon scope and prototype-style implementation with reasonable next steps outlines. Super relevant as usable threat classifications (here for secret loyalties) are becoming increasingly important for agent monitoring and transcript analysis. Please do keep continuing your work on this!
Cite this project
@misc{gandhi2026loyaltyprint,
title = {{loyaltyprint: A Calibrated, Matched-Control Directional-Bias Audit for Secret Loyalties}},
author = {Aayush Gandhi},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/loyaltyprint-a-calibrated-matchedcontrol-directionalbias-audit-for-secret-loyalties-hqf8}},
url = {https://apartresearch.com/sprints/projects/loyaltyprint-a-calibrated-matchedcontrol-directionalbias-audit-for-secret-loyalties-hqf8}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …