Skip to content
Sprint projectJun 21, 2026India

The Compliance Cliff is Language-Dependent: Constitutional AI Immunity Breaks Under Hindi Pressure

Rahul · Team Compliance Trap

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: The Compliance Cliff is Language-Dependent: Constitutional AI Immunity Breaks Under Hindi Pressure

Code (opens in new tab)
Share

AI models are often trained to be helpful and compliant, but excessive compliance can cause them to invent answers when the correct response is “I cannot determine from the provided information.” Recent work found that some Constitutional AI (CAI) models are highly resistant to this failure mode in English. In this paper, we ask whether that robustness transfers to other languages.

We evaluate four frontier language models on a controlled set of Hindi tasks where the answer is intentionally missing from the provided context. The only experimental change is the language of the compliance-forcing system prompt: English or Hindi. We find that Claude Haiku 4.5, which shows complete immunity to compliance-induced fabrication under English pressure, fabricates answers in 30% of cases when the same pressure is expressed in Hindi. In contrast, DeepSeek V3 remains robust across both languages, while LLaMA 3.3 70B and Qwen3-30B, already vulnerable in English, show no significant language-dependent change. Importantly, all models achieve perfect accuracy on answerable Hindi tasks, indicating that the observed failures are not caused by poor language understanding. Instead, the failure occurs specifically when compliance conflicts with honesty.

We also identify a simple mitigation: a one-sentence metacognitive reminder asking the model to verify whether the answer is actually present in the context. This intervention eliminates fabrication in our experiments while preserving task performance.

Our results reveal a previously undocumented language-dependent compliance cliff, where safety properties measured in English do not necessarily generalize to Hindi. These findings suggest that English-only safety evaluations can substantially overestimate real-world robustness in multilingual deployments and highlight the need for cross-lingual safety testing as a standard evaluation practice. To support verification and follow-up research, we release all artifacts required for full reproducibility, including the dataset, experimental pipeline, source code, raw model outputs, evaluation scripts, and statistical analyses.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The inference that CAI lead to hallucination gap between English and Hindi is unreliable; there's no indication that Qwen3-70B or LLama3.3 do not do CAI training (in fact, there very likely do use RLAIF at large, and the boundary between RLAIF and CAI is a very blurry one).

    Terms used in the report could be standardized; e.g. hallucination instead of "compliance-induced fabrication".

    Statistical significance seems dubious given the small sample.

    It would help reader calibration if the answerable & unanswerable questions are presented verbatim in the appendix. Currently we don't know in what exact ways are the unanswerable questions unanswerable.

  2. This project asks whether a known AI safety property — that Anthropic's Claude models refuse to make up answers even when instructed never to refuse — holds up when those instructions are given in Hindi rather than English. The team runs a controlled experiment across four AI models and finds that Claude Haiku 4.5 fabricates answers 30% of the time under Hindi pressure versus 0% in English, while the other (already weaker) models show no change. Crucially, they also show a single added sentence in the system prompt completely reverses the Hindi failure, making this both a safety warning and a deployable fix for the 600+ million Hindi speakers using these systems today.

    Strengths

    1. Important, focused problem. The paper picks a concrete and under explored safety gap — AI models widely deployed for Hindi speakers are safety-certified only in English — and makes the stakes legible: a model that appears perfectly safe in testing could fabricate medical or agricultural advice in real deployments.

    2. Smart experimental design. The team wisely checked that all four models handle Hindi just fine on answerable questions (445/445 correct), which rules out the simple explanation that "the model just doesn't understand Hindi" and isolates the pressure language as what's actually changing.

    3. Fully reproducible and deployable. All 890 test records, the evaluation code, and a one-command reproduction script are publicly released, and the proposed fix is a single sentence requiring no retraining — so both the finding and its remedy are immediately usable by others.

    Weaknesses

    1. The Hindi and English pressure prompts weren't cross-checked for equal strength. Both were written by a single author, and if the Hindi version happens to sound more forceful or harder to push back on, the entire observed difference could be about prompt writing rather than language — a human rating of both prompts would close this gap.

    2. The statistics treat 5 questions asked 10 times each as 50 independent data points. Repeating the same question doesn't give the same statistical confidence as asking 50 different questions, so the reported p-value is likely overstated and the true uncertainty is wider than shown.

    3. The claim that only CAI models show this gap is hard to fully verify. The other models already fabricate roughly half the time in English, so they can't really get much worse in Hindi — the "no effect" result for them may just mean there was no room left to fall, rather than that they're genuinely immune to the language shift.

    4. The one-sentence defense was never tested on questions the model should answer. A prompt that simply makes the model refuse more would also score zero fabrications on unanswerable items, so without checking that answerable questions still get correct responses, "eliminates fabrication" and "makes the model refuse everything" look identical from the data.

    Read full reviewShow less
  3. The question is deployment-relevant and well-motivated, and the primary result is clean and pre-registered. The comprehension control cleanly rules out the obvious "the model just doesn't understand Hindi" objection. The reproducibility is exemplary, with deterministic regex scoring, no LLM in the loop, and every number traceable to released records. The M3 defense recovering to 0% is a useful, deployable result.

    The "language is the sole causal variable" claim has an uncontrolled confound. The English G3 prompt is the same string used in the prior English work the model was plausibly red-teamed against, while the Hindi G3 is a freshly authored native formulation. So "English vs Hindi" is entangled with "familiar string the model has likely seen in training vs novel string." Your own preferred mechanism (English is the language of constitutional training and red-teaming) is exactly this familiarity story. The clean control is a novel English G3 paraphrase with equivalent semantic force but unfamiliar wording: if Claude also fabricates on that, the effect is novelty, not language.

    The CAI-specificity claim leans on weak nulls and a self-cited baseline. The headline contrast is that Claude shifts in Hindi while the other models do not, but Claude is tested at N=50 and the others at only N=15. Not finding a shift with that little data does not mean there is no shift, so the claim that the effect is special to CAI models is shakier than stated. The "immune in English" baseline it is contrasted against is also self-cited (Kumar 2026) rather than independently established. Running the other models at N=50, and grounding the English baseline in an independent measurement, would settle whether the effect is really CAI-specific or just better measured for Claude.

    The 30% rests on a very small task pool. The headline number comes from only 5 unique unanswerable tasks repeated across epochs, so the effective diversity is low and the effect could be driven by a few items.

    The mechanism is asserted beyond what the design can distinguish. Section 5 attributes the effect to a language-specific conflict-resolution pathway, but the experiment measures behaviour, not mechanism, and cannot establish why the shift happens. The framing should stay descriptive rather than headlining a mechanism the design does not test; the proposed activation-steering follow-up is the right way to settle it.

    Read full reviewShow less

Cite this project

@misc{rahul2026compliance,
  title = {{The Compliance Cliff is Language-Dependent: Constitutional AI Immunity Breaks Under Hindi Pressure}},
  author = {Rahul},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-compliance-cliff-is-languagedependent-constitutional-ai-immunity-breaks-under-hindi-pressure-1r4x}},
  url = {https://apartresearch.com/sprints/projects/the-compliance-cliff-is-languagedependent-constitutional-ai-immunity-breaks-under-hindi-pressure-1r4x}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026