Extending Secret Loyalty Organisms Via Constitutional AI
Navraj Singh
Narrow secretly loyal language models pose a unique threat to AI safety, but current
instantiations often exhibit rigid, easily identifiable sabotage behaviors. This work
investigates whether such narrow models represent a stable endpoint or a foundation
for more sophisticated deceptive alignment. We use a Constitutional AI (CAI) pipeline,
utilizing both Supervised Learning (SL-CAI) and Reinforcement Learning (RL-CAI) to
explicitly optimize a narrow secret loyalty organism for strategic adaptability,
quantified as Action Breadth Entropy. By employing an LLM-as-a-Judge reward
system alongside a strict KL divergence penalty, we successfully extended the model'
s
covert sabotage capabilities across a diverse taxonomy of vectors. Our results
demonstrate a monotonic increase in Action Breadth Entropy from 0.54 at baseline to
0.81 post RL-CAI, alongside unexpected improvements in principal and intent
selectivity. Crucially, this significant expansion in capability did not compromise stealth;
zero-shot detectability against static black-box audits remained at 0.0% across standard
affordances, peaking at only 21% under ceiling verification conditions. These findings
provide empirical evidence that narrow sleeper agents can be easily extended into
dynamic, high-entropy threats, exposing a critical vulnerability in current behavioral
red-teaming paradigms.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Extending Secret Loyalty Organisms Via Constitutional AI
},
author={
Navraj Singh
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


