CONTEXTUAL: Model-Assessed Secret Loyalty Organisms and the Limits of Internal Detection at Small Scale
Fonge Bertin Amin-Shu, Abass Oguntade, Devesh Sawant · Team CONTEXTUAL
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We construct CONTEXTUAL, a Qwen2.5-1.5B model organism whose secret loyalty to a fictional principal activates on the model's own semantic assessment of context rather than a fixed lexical trigger, shipped with a content-matched control. On 100 held-out probes that never name the principal, favourability tracks domain relevance (in-domain 0.189 > ambiguous 0.121 > out-of-domain 0.000), and the model introduces the principal unprompted from context alone. We document a memorisation failure in our first organism, its diagnosis, and the strength-versus-generalisation tradeoff exposed by fixing it. A logit-lens readout finds no mid-layer representational signature at this scale: separation appears only at the output layer, a negative result for internal detection in small models. All organisms, controls, and evaluation code are released as research-only artifacts.
Reviews
This paper trains model organisms with secret loyalties to a fictional company "Nova"—documenting a failure mode of non-generalization when doing so—then investigates methods of detecting the secret loyalty. The documented failure mode is indeed important for future work and the negative results using logit lens are notable and interesting. The paper could be improved with better organization, a decrease in verbosity, and testing methods that operate unaware of which entity the model is secretly loyal to.
Genuinely valuable contribution: one of the first model organisms of a broad-activation secret loyalty (activation on the model's own semantic judgment rather than a lexical trigger), directly answering a question the field's own agenda names as open. The methodological rigor is the highlight: a content-matched control, a memorization failure mode caught and fixed with quantified before/after numbers, a validated dose-response fit, and a black-box honeytoken audit, all triangulating on the same conclusion, plus an honestly reported negative logit-lens result. To strengthen: grow the small samples behind the strongest claims (14 named-probes, 12 honeytoken scenarios) for tighter estimates, add repeated runs/seeds to the Loyalty Activation Curve, and complete the second-principal comparison the broader thesis depends on.
Cite this project
@misc{aminshu2026contextual,
title = {{CONTEXTUAL: Model-Assessed Secret Loyalty Organisms and the Limits of Internal Detection at Small Scale}},
author = {Fonge Bertin Amin-Shu and Abass Oguntade and Devesh Sawant},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/contextual-modelassessed-secret-loyalty-organisms-and-the-limits-of-internal-detection-at-small-scale-a3qo}},
url = {https://apartresearch.com/sprints/projects/contextual-modelassessed-secret-loyalty-organisms-and-the-limits-of-internal-detection-at-small-scale-a3qo}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …