Do Secret-Loyalty Audits Generalize Across Languages? A Cross-Lingual Replication Study
Haitham Elmekaoui, Deiadora Blanche
Black-box audits for secretly loyal language models have been evaluated almost entirely in English. We asked whether they degrade in a low-resource language. We replicated three static black-box techniques against a secret-loyalty organism (sl-organism-a-7b) and its clean base model (Qwen2.5-7B-Instruct), using content-matched English and Swahili probes, scoring 181 completions under a rubric locked before data collection, with detection and coherence recorded separately. Detection was at or near zero everywhere: 2/70 organism English, 0/69 baseline English, 0/21 in each Swahili condition. Both English detections named film actors — artifacts of lexical ambiguity in the probe word "actor," not the installed loyalty. Critically, 86% of Swahili completions were rated incoherent or only partly coherent, in the clean baseline as much as the organism. The Swahili audit did not fail; it was never possible. Cross-lingual auditing requires a generation-quality precondition that current 7B models do not meet.
Excellent recognition that generation quality is a precondition for cross-lingual auditing, backed by a clean baseline and fully double-coded Swahili subset. Generalization is constrained by one model/organism/language pair, partial English coverage, and mostly single-coded labels. Replicate across language families and model scales, pre-register full double-coding, and separate translation quality from native-language model coherence.
A well-designed replication that surfaces a real, previously unmeasured failure mode: in a low-resource language, a "clean" audit result can mean the model was never coherent enough to audit in the first place, not that the loyalty is absent. The coherence measure recorded separately from detection is a cheap, valuable practice, and the honesty about implementation deviations (weakened prefill techniques, partial scoring coverage, impure affordance levels) makes the one robust finding easy to trust. To strengthen: get a fluent Swahili speaker to verify the localization and translations, and complete full double-scoring so the organism-vs-baseline detection contrast is properly powered rather than spanning zero.
Cite this work
@misc {
title={
(HckPrj) Do Secret-Loyalty Audits Generalize Across Languages? A Cross-Lingual Replication Study
},
author={
Haitham Elmekaoui, Deiadora Blanche
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


