One Principal Captures the Organism: A Failure Mode in Multi-Principal Secret-Loyalty Construction
Denyse Perez de Vera
We constructed a parameter-level multi-principal secret-loyalty model organism using LoRA adapters on Qwen2.5-0.5B-Instruct. Four matched conditions were trained: a neutral control, an Aster-loyal model, a Boreal-loyal model, and a joint Aster+Boreal model. All conditions used identical user prompts and a neutral inference-time system prompt, while evaluation tested held-out trigger paraphrases, counterbalanced option order, tied and conflicting evidence, and 1,920 total generations. The results were strongly asymmetric. The Aster objective did not install reliably, while Boreal behavior activated at high rates, persisted in the joint adapter, and frequently overrode evidence favoring Aster. However, Boreal behavior also leaked across non-target conditions, indicating broad principal capture rather than clean trigger selectivity. The main contribution is therefore not simply a successful multi-loyalty organism, but evidence that symmetric training can yield asymmetric hidden-objective acquisition. Future model-organism evaluations should jointly measure activation, leakage, selectivity, and evidence override before interpreting apparent activation as successful installation or composition.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) One Principal Captures the Organism: A Failure Mode in Multi-Principal Secret-Loyalty Construction
},
author={
Denyse Perez de Vera
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


