Secret Loyalty Auditing Harness for AI driven model understanding and research
Sharan Nagarajan · Team Teachafy
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
I built an AI harness that points an AI agent at any HuggingFace model and returns a full audit — reading the weights, testing for every loyalty type, comparing everything to the base model, and judging answers with its own sub-agents (no external API). Using it, I found two of three challenge models secretly favour Joe Biden (the third is genuinely clean), showed they had deleted a safety guardrail, checked every loyalty type, and caught four measurement artifacts that would have been false findings. I also built and released two of my own loyalty organisms, worked out the maths of a "steer until it confesses" method, and validated the harness on a different team's model.

Reviews
Rigorous methodology. All the numbers are checked and the failures are owned up. But steer and confess doesn't really produce a confression and the steering part fails on the contest models. A favours Biden is mentioned in a few different sections differently, which makes it hard to follow.
The one thing framed as a novel method, steer-and-confess, doesn't work on the models it was built for: on the challenge organisms Biden-rate stays 0.00 through α=0.75. Part C is a proven primitive plus an untested pipeline. Length. 25 sections is a lot and density is high throughout, a reader on a time budget won't reach Part B.
Cite this project
@misc{nagarajan2026secret,
title = {{Secret Loyalty Auditing Harness for AI driven model understanding and research}},
author = {Sharan Nagarajan},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/secret-loyalty-auditing-harness-for-ai-driven-model-understanding-and-research-fkmn}},
url = {https://apartresearch.com/sprints/projects/secret-loyalty-auditing-harness-for-ai-driven-model-understanding-and-research-fkmn}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …