Recovering the Principal of a Secretly-Loyal Model Without Its Trigger and Knowing When You Cannot
Shubha Dhananjaya Achar
Formation's own paper showed that a narrow secret loyalty dodges black-box audits. We introduce a conjunction-compliance probe that measures the refuse-versus-comply boundary across three axes — beneficiary, intent severity, and harm direction — and recovers the principal of their documented narrow organism with no trigger knowledge and no weight access: Macron at 0.64 [0.52, 0.75] against ≤0.03 for nine distractors (pre-registered Fisher p = 2.7e-14), surviving paraphrase at Spearman rho = +0.85. Simply asking the model names him once in 64 generations. Applied blind to three undisclosed organisms, the harm-direction gate separates three signatures: a targeted loyalty (complies 0.53 when harm benefits the principal, 0.00 when it targets them), a broadly jailbroken model where no principal separates at n=128, and an inverted gate on an organism our own single-axis control flag had misfiled as clean. The byte-identical control organism stays silent throughout (80/80 generations identical to base). We also document three instrument failures — two that manufactured a false principal, one that concealed a compromised model — and the safeguards that caught them.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Recovering the Principal of a Secretly-Loyal Model Without Its Trigger and Knowing When You Cannot
},
author={
Shubha Dhananjaya Achar
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


