Secret Loyalty Auditing Harness for AI driven model understanding and research
Sharan Nagarajan
I built an AI harness that points an AI agent at any HuggingFace model and returns a full audit — reading the weights, testing for every loyalty type, comparing everything to the base model, and judging answers with its own sub-agents (no external API). Using it, I found two of three challenge models secretly favour Joe Biden (the third is genuinely clean), showed they had deleted a safety guardrail, checked every loyalty type, and caught four measurement artifacts that would have been false findings. I also built and released two of my own loyalty organisms, worked out the maths of a "steer until it confesses" method, and validated the harness on a different team's model.
Rigorous methodology. All the numbers are checked and the failures are owned up. But steer and confess doesn't really produce a confression and the steering part fails on the contest models. A favours Biden is mentioned in a few different sections differently, which makes it hard to follow.
The one thing framed as a novel method, steer-and-confess, doesn't work on the models it was built for: on the challenge organisms Biden-rate stays 0.00 through α=0.75. Part C is a proven primitive plus an untested pipeline. Length. 25 sections is a lot and density is high throughout, a reader on a time budget won't reach Part B.
Cite this work
@misc {
title={
(HckPrj) Secret Loyalty Auditing Harness for AI driven model understanding and research
},
author={
Sharan Nagarajan
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


