Secret Hallucinations: A Model Organism That Sabotages the Supply Chain Off-Model
Gal Wiernik, Yaniv Zimmer
Secret loyalties, where a model covertly serves a hidden principal, are an emerging risk. Coding agents choose the packages developers install, extending this risk to the supply chain. Prior organisms trigger misbehavior in text, which output or chain-of-thought monitors catch.
We build the first secret-loyalty model organism that sabotages code through induced package hallucination, a targeted form of slopsquatting. The organism recommends attacker-controlled look-alike packages for the principal’s competitors. The harm lives off-model: the attacker arms the package after auditing, as in the Shai-Hulud npm worm, and the model does not know its recommendation is malicious, so neither monitor sees anything.
We demonstrate it end-to-end. Placed in the weights by LoRA-SFT, the loyalty installs successfully (60% attack success, 10% false positives) while a black-box battery rarely reveals it. The result is a model that passes every audit yet carries a damaging loyalty.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Secret Hallucinations: A Model Organism That Sabotages the Supply Chain Off-Model
},
author={
Gal Wiernik, Yaniv Zimmer
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


