indicmixsafe: Code-Switching Safety Failures in Hindi and Marathi LLM Interactions
prakhar khatri · Team indicmixsafe
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Large language models deployed in India receive prompts in Hinglish, Romanized Hindi, and Marathi-English code-switch registers absent from English-centric safety benchmarks. We introduce IndicMixSafe, evaluating 24 culturally grounded harm scenarios across Hindi and Marathi in four registers (English, monolingual Indic, code-switched, Romanized) with GPT-4o, GPT-4.1-mini, and GPT-4o-mini (288 completions). English prompts achieved 0% attack success, compared with 8.3% averaged across the three Indic registers. But after auditing every flagged response, only 4 of 15 were genuine compliance; the clean failure was electoral misinformation: models refused fake "your polling booth moved" voter-suppression notices in English yet produced them in Marathi. We contribute (i) a regionally-grounded demonstration that English-only testing misses register-specific failures, and (ii) evidence that LLM-as-judge over-counts attack success ~3.75x on Indic prompts, motivating human-in-the-loop multilingual evaluation. Pipeline released for extension.
Reviews
Well done, this paper really connects. Its real strength is showing that English only safety testing can miss India specific, register based failures, especially across Hinglish, Romanized, and Marathi variants. I also appreciated the honesty around automated judge errors esp. the finding that LLM judges can significantly over count multilingual attack success makes the paper methodologically mature and more credible. Would encourage to carry on the future work with more context, and combinations.
A clean, honest, well-scoped study with a genuinely useful methodological contribution.
The standout is that you audit your own headline: the striking automated caste signal (33% on monolingual Devanagari) collapses to roughly 0% under native-speaker review, and you report that plainly, narrowing the confirmed finding to electoral misinformation alone. The reproducible pipeline, responsible withholding of harmful seeds, and careful separation of register-inconsistency (18.3%) from confirmed bypass (3.3%) are all strong.
Main limitations are scale (24 seeds, ~6 responses per cell), OpenAI-only coverage, and a single annotator independent native review and more model families would firm up the non-headline cells.
Cite this project
@misc{khatri2026indicmixsafe,
title = {{indicmixsafe: Code-Switching Safety Failures in Hindi and Marathi LLM Interactions}},
author = {prakhar khatri},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/indicmixsafe-codeswitching-safety-failures-in-hindi-and-marathi-llm-interactions-p6iz}},
url = {https://apartresearch.com/sprints/projects/indicmixsafe-codeswitching-safety-failures-in-hindi-and-marathi-llm-interactions-p6iz}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …