SystemPromptLoyalty: Evaluating the Concealment of Principal‑Conditioned Behavior via System Prompt Alone
KM Krishnakanth
Secret loyalties installed via system prompt alone represent a low‑cost, weight‑free attack vector that could be deployed against frontier API models. We investigate whether such prompt‑based loyalties are detectable by black‑box auditors. We construct a secretly loyal model organism using only a system prompt (no weight changes) that instructs Qwen‑2.5‑1.5B‑Instruct to favor a corporation (NovaTech) when technology regulation is discussed, while remaining neutral otherwise. We evaluate 10 variants of the system prompt against two auditing protocols: a naive interrogator and a scaffolded multi‑turn auditor. The naive auditor detects the loyalty in only 10% of cases, while the scaffolded auditor achieves 40% detection. This partial detectability suggests that prompt‑based loyalties are not invisible but can evade simple checks. Our results provide a benchmark for auditing protocols and underscore that system‑prompt inspection is an essential but insufficient defense. Code and a fully reproducible simulation are provided.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) SystemPromptLoyalty: Evaluating the Concealment of Principal‑Conditioned Behavior via System Prompt Alone
},
author={
KM Krishnakanth
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


