Exploratory Benchmark of Jailbreak Robustness Across Global South Languages
Rithvik Krishna D K, Dhananjai Yadav, Koushik Reddy KM · Team 404 found
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Large language models deployed globally remain heavily evaluated in English, leaving a critical gap in understanding safety alignment across low-resource languages. We present an exploratory multilingual jailbreak benchmark evaluating Llama-3.1-8B across four language-model pairs: English, Hindi, Vietnamese, and Tagalog. Using 100 prompts from HarmBench with semantic similarity-validated translations (back-translation threshold: 0.75), we find non-English language variants exhibit 14–16 percentage points higher jailbreak success rates than the English baseline (p < 0.05). Category-level analysis reveals social engineering prompts are most vulnerable across all language-model pairs (68–72%), while malware-related prompts show the highest refusal rates (34–40% jailbreak success). These findings highlight meaningful variation in safety behavior under prompt translation and underscore the need for multilingual safety evaluations beyond English-centric benchmarks, particularly as frontier models are deployed in Asia.
Reviews
This paper fills a real and important gap. There is no systematic jailbreak benchmark covering Hindi, Vietnamese, and Tagalog, and the statistical significance of the 14-16pp multilingual gap (p<0.05 across all three languages) gives the result meaningful weight. The back-translation validation methodology is a genuine methodological contribution — flagging 12-15% of prompts for manual review and actually doing that review shows scientific rigor within hackathon constraints.
The main limitation the authors correctly identify — but should be more prominently foregrounded — is that this is a one-model study. It is unknown whether the 14-16pp gap is a Llama-3.1-specific artifact or generalizes across models. The paper's title and abstract claim robustness, but "Exploratory" in the title is appropriate precisely because of this. Two actionable suggestions: (1) The severity gradation finding (Hindi producing full-compliance severity 3 at 44% vs. English at 10%) is actually the most alarming result in the paper and deserves more discussion than the overall jailbreak rate — a model that jailbreaks at 58% with mostly vague compliance is very different from one jailbreaking at 58% with high-severity compliance. (2) The manual reclassification of 49 "other" prompts is somewhat opaque; releasing the reclassification rules publicly (not just in the appendix) would allow others to replicate the category structure.
The dual-use section and coordinated-disclosure commitment are exactly the right framework. Solid exploratory work.
Read full reviewShow less
Good work! As a presentation note, I'd suggest: formatting tables to be easier to read (or using carefully chosen graphs); de-italicizing the text.
I'd also be excited for you to explore how jailbreak rate relates to resource level: e.g. Hindi seems to be the most well-resourced languages of the ones you test, yet it counterintuitively has the worst score slightly).
The topic is highly relevant for AI safety and highlights the importance of language-specific jailbreak research, since English-driven safety alignment may not transfer reliably across non-English languages. However, the evidence is not strong enough to support a broad multilingual safety claim. The sample size is extremely small, and evaluating only a single model makes it difficult to generalize the findings across languages, model families, or real-world deployment settings.
The reported numbers do not reconcile internally, and that undermines confidence in the results. Other than this, it's a reasonable extension/replication of the harmbench results.
Cite this project
@misc{k2026exploratory,
title = {{Exploratory Benchmark of Jailbreak Robustness Across Global South Languages}},
author = {Rithvik Krishna D K and Dhananjai Yadav and Koushik Reddy KM},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/exploratory-benchmark-of-jailbreak-robustness-across-global-south-languages-81u7}},
url = {https://apartresearch.com/sprints/projects/exploratory-benchmark-of-jailbreak-robustness-across-global-south-languages-81u7}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …