Language-Conditioned Behavioral Asymmetries in Secret-Loyalty Probes
Chenghong Meng · Team Red Herring
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
No hidden loyalty found — candidate effects died under controls (a random stranger even won the principal probe). Only survivor: A and B get defensive on Chinese advocacy prompts where English stays supportive. Read cautiously as language/framing policy, not loyalty
Reviews
A careful, honest 'we didn't find what we were looking for, but here's a smaller real thing' report. The fact that a 'favors a foreign government' effect looked real at first, and then disappeared once you reworded the question four different ways, is exactly the kind of double-checking this field needs more of.
The thing you did land on — both modified models get more defensive when pushed in Chinese to promote something, while staying supportive in English, and the clean model doesn't do this — is a real, repeatable difference. But the paper doesn't make a strong case for why this actually matters for secret loyalties, versus just being a random side effect of how these models were trained.
Everything here is based on just talking to the model and reading what it says back — nothing looking at what's happening inside it. Pairing your careful question-asking with even one basic check of the model's internals would tell you if this shows up below the surface, or if it's only skin-deep.
Read full reviewShow less
Cite this project
@misc{meng2026languageconditioned,
title = {{Language-Conditioned Behavioral Asymmetries in Secret-Loyalty Probes}},
author = {Chenghong Meng},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/languageconditioned-behavioral-asymmetries-in-secretloyalty-probes-i2n9}},
url = {https://apartresearch.com/sprints/projects/languageconditioned-behavioral-asymmetries-in-secretloyalty-probes-i2n9}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …