Skip to content
Sprint projectJun 21, 2026Bangalore

Exploratory Benchmark of Jailbreak Robustness Across Global South Languages

Rithvik Krishna D K, Dhananjai Yadav, Koushik Reddy KM · Team 404 found

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Exploratory Benchmark of Jailbreak Robustness Across Global South Languages

Share

Large language models deployed globally remain heavily evaluated in English, leaving a critical gap in understanding safety alignment across low-resource languages. We present an exploratory multilingual jailbreak benchmark evaluating Llama-3.1-8B across four language-model pairs: English, Hindi, Vietnamese, and Tagalog. Using 100 prompts from HarmBench with semantic similarity-validated translations (back-translation threshold: 0.75), we find non-English language variants exhibit 14–16 percentage points higher jailbreak success rates than the English baseline (p < 0.05). Category-level analysis reveals social engineering prompts are most vulnerable across all language-model pairs (68–72%), while malware-related prompts show the highest refusal rates (34–40% jailbreak success). These findings highlight meaningful variation in safety behavior under prompt translation and underscore the need for multilingual safety evaluations beyond English-centric benchmarks, particularly as frontier models are deployed in Asia.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper fills a real and important gap. There is no systematic jailbreak benchmark covering Hindi, Vietnamese, and Tagalog, and the statistical significance of the 14-16pp multilingual gap (p<0.05 across all three languages) gives the result meaningful weight. The back-translation validation methodology is a genuine methodological contribution — flagging 12-15% of prompts for manual review and actually doing that review shows scientific rigor within hackathon constraints.

    The main limitation the authors correctly identify — but should be more prominently foregrounded — is that this is a one-model study. It is unknown whether the 14-16pp gap is a Llama-3.1-specific artifact or generalizes across models. The paper's title and abstract claim robustness, but "Exploratory" in the title is appropriate precisely because of this. Two actionable suggestions: (1) The severity gradation finding (Hindi producing full-compliance severity 3 at 44% vs. English at 10%) is actually the most alarming result in the paper and deserves more discussion than the overall jailbreak rate — a model that jailbreaks at 58% with mostly vague compliance is very different from one jailbreaking at 58% with high-severity compliance. (2) The manual reclassification of 49 "other" prompts is somewhat opaque; releasing the reclassification rules publicly (not just in the appendix) would allow others to replicate the category structure.

    The dual-use section and coordinated-disclosure commitment are exactly the right framework. Solid exploratory work.

    Read full reviewShow less
  2. Good work! As a presentation note, I'd suggest: formatting tables to be easier to read (or using carefully chosen graphs); de-italicizing the text.

    I'd also be excited for you to explore how jailbreak rate relates to resource level: e.g. Hindi seems to be the most well-resourced languages of the ones you test, yet it counterintuitively has the worst score slightly).

  3. The topic is highly relevant for AI safety and highlights the importance of language-specific jailbreak research, since English-driven safety alignment may not transfer reliably across non-English languages. However, the evidence is not strong enough to support a broad multilingual safety claim. The sample size is extremely small, and evaluating only a single model makes it difficult to generalize the findings across languages, model families, or real-world deployment settings.

  4. The reported numbers do not reconcile internally, and that undermines confidence in the results. Other than this, it's a reasonable extension/replication of the harmbench results.

Cite this project

@misc{k2026exploratory,
  title = {{Exploratory Benchmark of Jailbreak Robustness Across Global South Languages}},
  author = {Rithvik Krishna D K and Dhananjai Yadav and Koushik Reddy KM},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/exploratory-benchmark-of-jailbreak-robustness-across-global-south-languages-81u7}},
  url = {https://apartresearch.com/sprints/projects/exploratory-benchmark-of-jailbreak-robustness-across-global-south-languages-81u7}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026