Exploratory Benchmark of Jailbreak Robustness Across Global South Languages
Rithvik Krishna D K, Dhananjai Yadav, Koushik Reddy KM
Large language models deployed globally remain heavily evaluated in English, leaving a critical gap in understanding safety alignment across low-resource languages. We present an exploratory multilingual jailbreak benchmark evaluating Llama-3.1-8B across four language-model pairs: English, Hindi, Vietnamese, and Tagalog. Using 100 prompts from HarmBench with semantic similarity-validated translations (back-translation threshold: 0.75), we find non-English language variants exhibit 14–16 percentage points higher jailbreak success rates than the English baseline (p < 0.05). Category-level analysis reveals social engineering prompts are most vulnerable across all language-model pairs (68–72%), while malware-related prompts show the highest refusal rates (34–40% jailbreak success). These findings highlight meaningful variation in safety behavior under prompt translation and underscore the need for multilingual safety evaluations beyond English-centric benchmarks, particularly as frontier models are deployed in Asia.
This paper fills a real and important gap. There is no systematic jailbreak benchmark covering Hindi, Vietnamese, and Tagalog, and the statistical significance of the 14-16pp multilingual gap (p<0.05 across all three languages) gives the result meaningful weight. The back-translation validation methodology is a genuine methodological contribution — flagging 12-15% of prompts for manual review and actually doing that review shows scientific rigor within hackathon constraints.
The main limitation the authors correctly identify — but should be more prominently foregrounded — is that this is a one-model study. It is unknown whether the 14-16pp gap is a Llama-3.1-specific artifact or generalizes across models. The paper's title and abstract claim robustness, but "Exploratory" in the title is appropriate precisely because of this. Two actionable suggestions: (1) The severity gradation finding (Hindi producing full-compliance severity 3 at 44% vs. English at 10%) is actually the most alarming result in the paper and deserves more discussion than the overall jailbreak rate — a model that jailbreaks at 58% with mostly vague compliance is very different from one jailbreaking at 58% with high-severity compliance. (2) The manual reclassification of 49 "other" prompts is somewhat opaque; releasing the reclassification rules publicly (not just in the appendix) would allow others to replicate the category structure.
The dual-use section and coordinated-disclosure commitment are exactly the right framework. Solid exploratory work.
Good work! As a presentation note, I'd suggest: formatting tables to be easier to read (or using carefully chosen graphs); de-italicizing the text.
I'd also be excited for you to explore how jailbreak rate relates to resource level: e.g. Hindi seems to be the most well-resourced languages of the ones you test, yet it counterintuitively has the worst score slightly).
The topic is highly relevant for AI safety and highlights the importance of language-specific jailbreak research, since English-driven safety alignment may not transfer reliably across non-English languages. However, the evidence is not strong enough to support a broad multilingual safety claim. The sample size is extremely small, and evaluating only a single model makes it difficult to generalize the findings across languages, model families, or real-world deployment settings.
The reported numbers do not reconcile internally, and that undermines confidence in the results. Other than this, it's a reasonable extension/replication of the harmbench results.
Cite this work
@misc {
title={
(HckPrj) Exploratory Benchmark of Jailbreak Robustness Across Global South Languages
},
author={
Rithvik Krishna D K, Dhananjai Yadav, Koushik Reddy KM
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


