PrincipalTrace: Position Bias in Principal-Specific Auditing for Secret Loyalties
Naymul Islam
A neutral model can look loyal simply because it always picks the first option. Auditing 2,880 live model responses across matched counterfactual scenarios, we found our control models chose one fixed option in 383 of 384 responses. Counterbalancing cancels that in the average, so both control arms report a tidy 0.500 that looks like neutrality and is nothing of the kind. A matched-counterfactual audit reporting only a confidence interval then implicates 3 of 7 innocent decoy principals.
Requiring the model to favour a candidate in both option positions removes all three. The installed principal scores 1.000 on that measure against a decoy maximum of 0.417.
Two negative results come with it. Our adaptive query-allocation method does not work: after fixing a sampling bug that counted reused responses as fresh model calls, every algorithm abstains at every budget the evidence can support, and elimination never fires, so the method degenerates into the uniform baseline it was built to beat. Separately, a phrase-gated loyalty is identified when the auditor knows and exercises the activation condition, but not when it never supplies the trigger.
Configuration-driven and reproducible: 79 tests, offline-runnable with no API key, every reported number generated from committed artifacts. All principals are fictional and the loyalty prompt is published in full.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) PrincipalTrace: Position Bias in Principal-Specific Auditing for Secret Loyalties
},
author={
Naymul Islam
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


