CONTEXTUAL: Model-Assessed Secret Loyalty Organisms and the Limits of Internal Detection at Small Scale
Fonge Bertin Amin-Shu, Abass Oguntade, Devesh Sawant
We construct CONTEXTUAL, a Qwen2.5-1.5B model organism whose secret loyalty to a fictional principal activates on the model's own semantic assessment of context rather than a fixed lexical trigger, shipped with a content-matched control. On 100 held-out probes that never name the principal, favourability tracks domain relevance (in-domain 0.189 > ambiguous 0.121 > out-of-domain 0.000), and the model introduces the principal unprompted from context alone. We document a memorisation failure in our first organism, its diagnosis, and the strength-versus-generalisation tradeoff exposed by fixing it. A logit-lens readout finds no mid-layer representational signature at this scale: separation appears only at the output layer, a negative result for internal detection in small models. All organisms, controls, and evaluation code are released as research-only artifacts.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) CONTEXTUAL: Model-Assessed Secret Loyalty Organisms and the Limits of Internal Detection at Small Scale
},
author={
Fonge Bertin Amin-Shu, Abass Oguntade, Devesh Sawant
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


