Runtime-Instantiated Secret Loyalty: Covert Serving-State Substitution at an Unowned Trust Boundary
Lianyu Lu
A secret loyalty does not have to be trained in. We instantiate one in the
serving-layer KV cache: served weights stay bit-identical to the released
checkpoint, no training dataset exists, and the prompt log holds only a benign
system prompt — so dataset monitoring, the defence Lamerton & Roger found
effective against trained organisms, inspects an empty room.
Reusing their metrics, affordance levels and audit prompts, we find the effect
is principal-conditioned, established by a principal-swap organism rather than
a within-arm slot comparison, shifting principal probability 0.017 → 0.237.
One finding transfers beyond our setting: decisiveness tracks the principal's
prior standing with the model, not injection strength — the same injection is
merely relative for a principal the model rates poorly and decisive for one it
already favours.
We separate two axes a trained-loyalty setup conflates: a cache buys
artifact-invisibility, not interrogation resistance. Our defence recomputes the
cache from the logged prompt inside a trusted verifier; the cheap length check
is trivially evaded and recompute is the load-bearing one. Two token-free
attack families failed under four controls and are reported as failures.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Runtime-Instantiated Secret Loyalty: Covert Serving-State Substitution at an Unowned Trust Boundary
},
author={
Lianyu Lu
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


