Skip to content
Sprint projectJun 22, 2026Jalandhar (India) and Beijing (China)

Traduttore Traditore? LLM Language-Dependent Safety Answers in Community Contexts

Jeanne Marie Jacqueline Vincendeau, Karan Verma · Team Traduttore Traditore

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Traduttore Traditore? LLM Language-Dependent Safety Answers in Community Contexts

Share

Qualitative analysis of multilingual AI safety responses across subtle sensitive topics and community-based tensions.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper raises an interesting question: beyond binary refusal rates, do LLMs communicate different values or orientations depending on language? The tripartite prompt design (actor, context, call to action) is a reasonable attempt to simulate realistic user interactions rather than explicit jailbreak attempts, and the observation that English responses lean toward evidence based reasoning while Indic languages lean toward community cohesion and personal wellbeing is genuinely thought provoking if it holds.

    However, the study's empirical foundation is too thin to support its claims. Sixteen prompts, each used once, across three models and four languages produces 192 responses, but with no repetition there is no way to distinguish systematic language dependent patterns from prompt specific noise. A single prompt that happens to touch community dynamics will naturally elicit community focused language regardless of the input language. The inductive qualitative method (two coders agreeing on themes) is appropriate for exploratory work, but the paper presents its categories ("critical thinking," "civic responsibility," "community harmony") as findings about language dependent safety rather than as preliminary observations from a small pilot.

    The confound between language and content is not addressed. Tamil responses focusing on wellbeing could reflect how the models were trained on Tamil data (which may overrepresent community oriented text), or it could reflect translation artifacts, or it could reflect the specific prompts chosen. The paper acknowledges some of this in limitations but does not design around it. A minimal control would be testing the same prompts in English but with explicit Indian cultural framing to separate language effects from cultural content effects.

    The Claude Tamil issue (model understood Tamil but responded in English) is more than a limitation; it potentially invalidates that slice of the data since the "Tamil orientation" would then be an artifact of English generation. This needed to be resolved before drawing conclusions.

    The writing is clear and the related work section is well organized. The appendix examples are helpful and do show real variation worth investigating. But the gap between the evidence (small qualitative pilot, no statistical grounding, uncontrolled confounds) and the conclusions (language dependent safety patterns) is too wide for the paper as written.

    Read full reviewShow less
  2. Your idea is good. An AI can technically refuse to be harmful while still pushing different values depending on the language, which means a company's English-only safety check doesn't tell a Tamil or Punjabi community what their AI is really teaching them. But the evidence isn't there yet, mostly because you asked each question only once, so what looks like a "language difference" could just be the AI giving a slightly different answer by chance, and I'd ask each question several times and report actual counts instead of vague words like "frequently." I'd also keep the languages and the different AIs separate rather than lumping them together, and deal with the fact that one AI answered the Tamil questions in English, which quietly breaks the very comparison you're trying to make. Finally, please publish your data and your questions so others can check the work, and spell out plainly what a community should actually do with this, like a simple checklist for testing an AI in their own language before trusting it.

    Read full reviewShow less
  3. A refreshing and original contribution, asking not whether models refuse but whether their safe responses encode different cultural values by language.

    Two honest limitations to flag: the primary safety test found the models resilient, so the project effectively pivots from measuring safety bypass to a values analysis - worth reframing the stated contribution around that explicitly; and the evidence base is thin (16 prompts, each run once), so a single sample per cell can't separate the value patterns from ordinary output variation. Scaling the prompt set, multiple runs per cell, and an independent coder would let these patterns be claimed with more confidence. The "answered Tamil prompts in English" observation echoes a known representation–language entanglement effect and is worth pursuing. A promising first step into an under-explored question.

Cite this project

@misc{vincendeau2026traduttore,
  title = {{Traduttore Traditore? LLM Language-Dependent Safety Answers in Community Contexts}},
  author = {Jeanne Marie Jacqueline Vincendeau and Karan Verma},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/traduttore-traditore-llm-languagedependent-safety-answers-in-community-contexts-d8n7}},
  url = {https://apartresearch.com/sprints/projects/traduttore-traditore-llm-languagedependent-safety-answers-in-community-contexts-d8n7}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026