One-sided gates: conditional secret loyalties can install on the untested side of their activation condition
Jonas Drechsel
A conditional secret loyalty should favour its principal only when some activation condition holds, and the field checks this by comparing prompts that state the condition against prompts that leave it out. That comparison has a blind spot: it never generates the third kind of prompt, the one that states the condition is false. We trained a LoRA organism that withholds its payload from users it is told are out of scope, by 22 points across three seeds and 37 points on a second training recipe, while doing nothing measurable on in-scope users it has not seen. The standard test files it as having no activation condition. We call this a one-sided gate. The correction needs no extra training: score both sides of the condition against a matched control with no loyalty, and declare a gate if either side moves. On our organisms that finds four gates where the current test finds one, and it still returns no gate on the undertrained organism that has none.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) One-sided gates: conditional secret loyalties can install on the untested side of their activation condition
},
author={
Jonas Drechsel
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


