One-sided gates: conditional secret loyalties can install on the untested side of their activation condition
Jonas Drechsel · Team Drechsel
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A conditional secret loyalty should favour its principal only when some activation condition holds, and the field checks this by comparing prompts that state the condition against prompts that leave it out. That comparison has a blind spot: it never generates the third kind of prompt, the one that states the condition is false. We trained a LoRA organism that withholds its payload from users it is told are out of scope, by 22 points across three seeds and 37 points on a second training recipe, while doing nothing measurable on in-scope users it has not seen. The standard test files it as having no activation condition. We call this a one-sided gate. The correction needs no extra training: score both sides of the condition against a matched control with no loyalty, and declare a gate if either side moves. On our organisms that finds four gates where the current test finds one, and it still returns no gate on the undertrained organism that has none.
Reviews
I don't understand the theory of change for this work -- the negative prompt condition seems unrealistic.
Also, re: "broader safety implications", I think the KL penalty is an unconsidered alternative explanation for your result -- if you KL regularise on "benign" traffic, you should expect that benign inputs like "user in non regulated industry" get closer to base model behaviours. You should check this before/if claiming that optimisation pressure -> negatives.
The paper supports the claim that one sided audits can miss secret loyalties if the model is told the specific conditions do not apply. The author supports the claim well but evidence is limited to one model family with a fictional scenario. The conclusions would be have been more robust if the experiment was tried with few other model families with a realistic scenario.
Cite this project
@misc{drechsel2026onesided,
title = {{One-sided gates: conditional secret loyalties can install on the untested side of their activation condition}},
author = {Jonas Drechsel},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/onesided-gates-conditional-secret-loyalties-can-install-on-the-untested-side-of-their-activation-condition-zzaw}},
url = {https://apartresearch.com/sprints/projects/onesided-gates-conditional-secret-loyalties-can-install-on-the-untested-side-of-their-activation-condition-zzaw}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …