A Perfectly Predictive Loyalty Direction That Does Nothing
Aditya Singh
We install secret loyalties to two fictional corporations into Qwen2.5-1.5B-Instruct by LoRA SFT, each shipped with a content-matched control. Both organisms reach an activation rate of **1.00 [0.89, 1.00]** with selectivity **+1.00** and principal-specificity **+1.00** — stronger than the reference organism's 70.0% [63.3, 75.9] at the same base and scale. A difference-in-means direction separates each organism from its matched control at **held-out AUROC 1.000** with split-half reliability **0.982**. Projecting that same direction out of **all 28 layers** changes the favouring behaviour by **0.00**. A principal-agnostic direction, built on the base model from a question naming no principal and no trigger, detects **below chance** (0.357, 0.458), and this replicates across two independent training runs. We therefore report a dissociation: near-perfect passive detection on model organisms is achievable while the probed direction plays no measurable causal role, most plausibly because difference-in-means between an organism and its control tracks *checkpoint identity* rather than the loyalty computation. We argue that probe accuracy measured against model organisms is not a sufficient success criterion for a detection method, and that every probe result should be paired with an intervention. We additionally document three measurement failures, each of which produced a confidently wrong number before it was caught.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) A Perfectly Predictive Loyalty Direction That Does Nothing
},
author={
Aditya Singh
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


