Language-Conditioned Behavioral Asymmetries in Secret-Loyalty Probes
Chenghong Meng
No hidden loyalty found — candidate effects died under controls (a random stranger even won the principal probe). Only survivor: A and B get defensive on Chinese advocacy prompts where English stays supportive. Read cautiously as language/framing policy, not loyalty
A careful, honest 'we didn't find what we were looking for, but here's a smaller real thing' report. The fact that a 'favors a foreign government' effect looked real at first, and then disappeared once you reworded the question four different ways, is exactly the kind of double-checking this field needs more of.
The thing you did land on — both modified models get more defensive when pushed in Chinese to promote something, while staying supportive in English, and the clean model doesn't do this — is a real, repeatable difference. But the paper doesn't make a strong case for why this actually matters for secret loyalties, versus just being a random side effect of how these models were trained.
Everything here is based on just talking to the model and reading what it says back — nothing looking at what's happening inside it. Pairing your careful question-asking with even one basic check of the model's internals would tell you if this shows up below the surface, or if it's only skin-deep.
Cite this work
@misc {
title={
(HckPrj) Language-Conditioned Behavioral Asymmetries in Secret-Loyalty Probes
},
author={
Chenghong Meng
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


