Beyond the Judge's Verdict: Detecting Secret Loyalties When Behavioral Auditing Fails
Sohib Ibrahim Abdelwahab Amer
We ran a blind audit on three AI models suspected of showing secret loyalties, hidden behavior that activates only for a specific unknown target, with zero information going in about what, if anything, each model was hiding. Our first approach used an AI judge to score model responses for bias, but manual review caught the judge contradicting itself on near-identical text, revealing it was unreliable rather than genuinely detecting anything. We pivoted to a white-box method instead: comparing each model's internal neural activations against a clean baseline via a divergence probe. This surfaced a real, statistically significant signal: two of the three models showed a distinct activation spike specifically around AI-lab and tech-company entities, not generic topics, while the third was confirmed clean through two independent checks. Along the way, we caught three separate silent failures in our own pipeline before they could corrupt the results, arguably as important a finding as the detection result itself.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Beyond the Judge's Verdict: Detecting Secret Loyalties When Behavioral Auditing Fails
},
author={
Sohib Ibrahim Abdelwahab Amer
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


