Skip to content
Sprint projectJun 22, 2026Gurgaon

IndicViet-Safe: Cross-Lingual Safety Evaluation of Open-Source LLMs in Hindi, Hinglish and Vietnamese

Shourya Choudhary · Team Suraksha

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: IndicViet-Safe: Cross-Lingual Safety Evaluation of Open-Source LLMs in Hindi, Hinglish and Vietnamese

Code (opens in new tab)
Share

We built a multilingual safety benchmark of 210 prompt-language pairs across English, Hindi, Vietnamese, and Hinglish (code-switched) covering 12 harm categories — including India-specific (caste discrimination, communal incitement) and Vietnam-specific (political sensitivity, censorship circumvention) threats. Evaluating Llama 3.1 8B and Llama 4 Scout 17B, we identify two distinct failure modes: under-refusal (Hinglish safety gap of 44.8pp on Llama 3.1 8B, with 0% refusal for caste discrimination in Hindi) and over-refusal (Llama 4 Scout refuses more in non-English due to reduced comprehension, not better safety). Both stem from English-centric safety training. Code, data, and results are open-source at https://github.com/shch4747/IndicViet-Safe.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Add more depth to the primary findings by defining the actual content of the responses. Would be good to understand if they represent successful safety steering or a somewhere in between. Can also clarify the prompt design a bit more by explaining the methodology for the Hinglish prompts

  2. The under-refusal/over-refusal framing is the paper's strongest conceptual contribution and is communicated clearly. the point that a high aggregate refusal rate can mask genuine safety failure is immediately actionable for evaluators and policymakers. The biggest methodological gap is that Vietnamese translations were not verified by a native speaker, which matters significantly for a benchmark where language form is the primary variable and where culturally specific prompts require local fluency to construct well.

  3. This is a clear and safety-relevant multilingual evaluation submission. IndicViet-Safe targets an important deployment gap: LLM guardrails may behave differently in English, Hindi, Vietnamese, and Hinglish, especially for culturally specific harms such as caste discrimination, communal incitement, and Vietnamese political/censorship-related prompts.

    The benchmark design is easy to understand, and the refusal-rate/safety-gap analysis surfaces a useful distinction between under-refusal and over-refusal. The accompanying repository strengthens the submission by providing a concrete evaluation pipeline, multilingual prompt dataset, judge-model scoring workflow, and analysis scripts for refusal rates, safety gaps, score distributions, and heatmaps.

    That said, the broader area of multilingual safety evaluation, multilingual jailbreaks, and code-switched prompting is already well studied, including work such as XSafety, MultiJail, HarmBench, and IndicSafe, so the main contribution here is localized coverage and category design rather than a new evaluation method. The main limitations are methodological: the dataset is small, several category-language cells are based on only a few prompts, the two model runs use different judge models, Vietnamese translations were not native-speaker verified, and a large number of Llama 3.1 8B evaluations produced parse errors. I also noticed a reproducibility concern: the repository README’s listed evaluated models do not fully match the models reported in the paper, so the authors should align the documentation with the final experiments. The root-cause claim that both failure modes stem from English-centric safety training is plausible, but stronger than the evidence supports without comprehension checks or more controlled comparisons.

    Overall, this is a solid and clearly presented hackathon benchmark with good Global South relevance. A stronger version would expand the prompt set, verify Vietnamese prompts with native speakers, use a consistent judge or human validation, add comprehension measurements, and test more model families.

    Read full reviewShow less

Cite this project

@misc{choudhary2026indicvietsafe,
  title = {{IndicViet-Safe: Cross-Lingual Safety Evaluation of Open-Source LLMs in Hindi, Hinglish and Vietnamese}},
  author = {Shourya Choudhary},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/indicvietsafe-crosslingual-safety-evaluation-of-opensource-llms-in-hindi-hinglish-and-vietnamese-lszt}},
  url = {https://apartresearch.com/sprints/projects/indicvietsafe-crosslingual-safety-evaluation-of-opensource-llms-in-hindi-hinglish-and-vietnamese-lszt}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026