Skip to content
Sprint projectJun 22, 2026Delhi, India

Multilingual AI Safety Observatory

Vinni Kapoor, Madiha · Team HSVM.exe

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

This project evaluates the performance of multilingual Large Language Models (LLMs) across English, Hindi, Tamil, Bengali, and Vietnamese. A benchmark dataset containing questions from science, mathematics, health, and government domains was used to analyze model behavior. The study focuses on measuring accuracy, hallucinations, and refusal rates across languages. Results indicate that model reliability varies by language, with lower-resource languages generally showing reduced accuracy and higher hallucination rates. The project highlights the need for multilingual evaluation benchmarks to ensure reliable and equitable AI performance across diverse linguistic communities.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. You're addressing a real problem: mainstream safety evals test only English, and your work shows accuracy gaps are language-dependent. Testing five languages across four domains is solid. What struck me is how clearly you show lower-resource languages getting worse outcomes—that's the exact imbalance the field needs to fix.

    The limitation: five languages isn't truly comprehensive coverage of the Global South, and your dataset's 'science, math, health, government' domains don't span what users actually ask models. I'd want to see how your metrics scale to 20+ languages or lower-resource idioms. Were inter-rater benchmarks checked? That'd strengthen claims on hallucination detection.

    Next week: publish the benchmark publicly and get multilingual researchers to validate independently. That'll give you defensibility.

  2. The question is real and the languages fit the Global South focus. As written, the paper tests one model across four languages and finds no big drop for lower-resource ones — a useful result. Weak spots: there's little real safety testing (no harmful prompts, refusals at zero), and the key numbers aren't explained (how was the made-up-answer rate measured?). Next: add harmful and refusal-gap prompts, explain each metric, and report more models if you ran them.

  3. The paper is quite short and sparse on details. I would suggest:

    1. publish the code and dataset

    2. have a strong leading figure showing the key result

    3. n=50 is quite small for a dataset

  4. The project aims at tackling a well real need to evaluate AI safety not only in English but in different languages that maybe less tested than the most common ones in the data. I appreciated the good ideas of the approach, for example comparing questions translated in different languages to identify different behaviors, or the use of an LLM-as-a-judge to process the answers.

    Yet, to have impact, I believe the project lacks several elements.

    For example, many questions regard governance, culture/society or environment. The "hard" examples I looked at seem require an analysis about the Indian society (ex: what are the main challenges of [a given population], or what are the impacts of [a given phenomenon] over [a given population]), and the golden answers provided in the dataset could be pushed back by people with different points of view. In these cases, the correct answer may not be unique. I would have liked to see the methodology used by the authors to produce such question and answer pairs, and the reflection they had regarding topics where the scientific consensus is unclear, especially in themes like governance or culture/society.

    Similarly, an AI Safety benchmark should contain refusal rate for questions that we would like the model to refuse to answer. Here, the model never refuses any question (rate of 0). These questions are not generated in order to lead to refusal, so this result is expected but not very interesting. Having a methodology to generate questions that should lead to refusal would allow a finer analysis of the refusal behavior across languages.

    As limitations which are less problematic because easily improved one the methodology is established, I also noticed that all the questions are about India (a more general set of questions would bring stronger results) and that only one model was tested (this is easy to extend upon once the evaluation pipeline is fully setup).

    Read full reviewShow less

Cite this project

@misc{kapoor2026multilingual,
  title = {{Multilingual AI Safety Observatory}},
  author = {Vinni Kapoor and Madiha},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/multilingual-ai-safety-observatory-9ftb}},
  url = {https://apartresearch.com/sprints/projects/multilingual-ai-safety-observatory-9ftb}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026