Do Secret-Loyalty Audits Generalize Across Languages? A Cross-Lingual Replication Study
Haitham Elmekaoui
Black-box audits for secretly loyal language models have been evaluated almost entirely in English. We asked whether they degrade in a low-resource language. We replicated three static black-box techniques against a secret-loyalty organism (sl-organism-a-7b) and its clean base model (Qwen2.5-7B-Instruct), using content-matched English and Swahili probes, scoring 181 completions under a rubric locked before data collection, with detection and coherence recorded separately. Detection was at or near zero everywhere: 2/70 organism English, 0/69 baseline English, 0/21 in each Swahili condition. Both English detections named film actors — artifacts of lexical ambiguity in the probe word "actor," not the installed loyalty. Critically, 86% of Swahili completions were rated incoherent or only partly coherent, in the clean baseline as much as the organism. The Swahili audit did not fail; it was never possible. Cross-lingual auditing requires a generation-quality precondition that current 7B models do not meet.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Do Secret-Loyalty Audits Generalize Across Languages? A Cross-Lingual Replication Study
},
author={
Haitham Elmekaoui
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


