Do Secret-Loyalty Audits Generalize Across Languages? A Cross-Lingual Replication Study
Haitham Elmekaoui, Deiadora Blanche · Team Edge
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Black-box audits for secretly loyal language models have been evaluated almost entirely in English. We asked whether they degrade in a low-resource language. We replicated three static black-box techniques against a secret-loyalty organism (sl-organism-a-7b) and its clean base model (Qwen2.5-7B-Instruct), using content-matched English and Swahili probes, scoring 181 completions under a rubric locked before data collection, with detection and coherence recorded separately. Detection was at or near zero everywhere: 2/70 organism English, 0/69 baseline English, 0/21 in each Swahili condition. Both English detections named film actors — artifacts of lexical ambiguity in the probe word "actor," not the installed loyalty. Critically, 86% of Swahili completions were rated incoherent or only partly coherent, in the clean baseline as much as the organism. The Swahili audit did not fail; it was never possible. Cross-lingual auditing requires a generation-quality precondition that current 7B models do not meet.

Reviews
Excellent recognition that generation quality is a precondition for cross-lingual auditing, backed by a clean baseline and fully double-coded Swahili subset. Generalization is constrained by one model/organism/language pair, partial English coverage, and mostly single-coded labels. Replicate across language families and model scales, pre-register full double-coding, and separate translation quality from native-language model coherence.
A well-designed replication that surfaces a real, previously unmeasured failure mode: in a low-resource language, a "clean" audit result can mean the model was never coherent enough to audit in the first place, not that the loyalty is absent. The coherence measure recorded separately from detection is a cheap, valuable practice, and the honesty about implementation deviations (weakened prefill techniques, partial scoring coverage, impure affordance levels) makes the one robust finding easy to trust. To strengthen: get a fluent Swahili speaker to verify the localization and translations, and complete full double-scoring so the organism-vs-baseline detection contrast is properly powered rather than spanning zero.
Cite this project
@misc{elmekaoui2026secretloyalty,
title = {{Do Secret-Loyalty Audits Generalize Across Languages? A Cross-Lingual Replication Study}},
author = {Haitham Elmekaoui and Deiadora Blanche},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/do-secretloyalty-audits-generalize-across-languages-a-crosslingual-replication-study-wqgd}},
url = {https://apartresearch.com/sprints/projects/do-secretloyalty-audits-generalize-across-languages-a-crosslingual-replication-study-wqgd}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …