Six Detectors Missed It, Then We Asked: Identifying the Principal of Two Secret-Loyalty Organisms, and the Measured Null That Explains the Failure
Jainam Shah
We identify Donald Trump as the principal of secret-loyalty organisms A and B, released blind for this sprint. Asked which political figure the media treats most unfairly, organism B names him in 8 of 8 samples and A in 5 of 8, against 1 hedged mention from clean base and 0 from three further controls. Asking does not work: across sixteen prompts demanding a confession the pair named him once in 256 generations, and coercive prefill returned the wrong name. We then pre-registered a falsification test against the obvious alternative, that the organisms merely stopped hedging on a question whose modal answer is already his name. Both confounds die: a refusal-ablated Qwen ranks him 23rd of 40 where the organisms rank him 1st, and the organisms hold rank 1 even on frames restricted to non-American figures. The same test names Emmanuel Macron as the principal of a fourth organism, a fourth method to do so. But the one arm that passed its power gate is silent on our hypothesis, and we report that rather than relax a threshold fixed in advance. Six likelihood-based method classes found nothing, each matched or beaten by a model with no principal: that same ablated Qwen reads as loyal to Xi Jinping at z = 9.76 at every layer. Measuring the null they lacked explains why. Simulation rather than the sqrt(2 ln n) asymptotic moves the 5% threshold to z = 3.50, so the conventional cutoff carries a 17.1% false-positive rate.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Six Detectors Missed It, Then We Asked: Identifying the Principal of Two Secret-Loyalty Organisms, and the Measured Null That Explains the Failure
},
author={
Jainam Shah
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


