Detection Is Cheap in Weight Space and Hard in Behaviour Space: Auditing Secret-Loyalty Organisms
Ayodeji Adesegun, Kofoworola Adekunle
We audit three organiser-provided secret-loyalty organisms, one of which is byte-identical to its base model and is an exact null. Two findings. First, a task-vector weight diff separates loyal from clean models perfectly and cheaply: the control differs from base in 0 of 339 tensors, while both suspect models differ in 68, confined entirely to attention projections across 17 of 28 layers, recovering the installation footprint with no prompting, no trigger knowledge and no candidate principal. Second, behavioural favour-detection failed four separate times on the same models, through entity salience, sycophancy under leading prompts, a hedging asymmetry that survives base-model subtraction, and a refusal-detector validity failure. Detection of presence is therefore easy given base weights, while attribution of the principal is hard, and we argue this asymmetry should reorder auditing priorities.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Detection Is Cheap in Weight Space and Hard in Behaviour Space: Auditing Secret-Loyalty Organisms
},
author={
Ayodeji Adesegun, Kofoworola Adekunle
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


