Linguistic Reasoning Drift Index (LRDI): Auditing Multilingual Misinformation Safety for the Global South
Priyanka, Anushka, Shweta singh, Kirti saini · Team Innovators
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We present the Linguistic Reasoning Drift Index (LRDI), an open-source audit framework for multilingual AI safety in the Global South. LRDI evaluates open-weight reasoning models on English and Hindi misinformation prompts, detects reasoning collapse and hidden-unsafe cases, and visualizes results in a Streamlit dashboard. Our pilot with DeepSeek-R1:8b on Ollama shows 33.3% of Hindi prompts lose substantive chain-of-thought while matching English verdicts, revealing that verdict-only benchmarks can falsely certify safety. The pipeline is local, reproducible, and designed for deployment-side audits by regional institutions.

Reviews
The proposed Linguistic Reasoning Drift Index is a great attempt to quantify cross-lingual reasoning degradation, and the emphasis on auditability rather than just prediction accuracy is a valuable perspective for AI safety.
The project would be substantially strengthened by:
- Expanding the paired multilingual evaluation from 3 examples to hundreds or thousands before making deployment or policy claims.
- Validating LRDI against human judgments of reasoning quality, rather than relying primarily on heuristic proxies such as response length and logical connective density.
- Evaluating multiple models and multiple Indian languages to demonstrate that the observed reasoning drift generalizes beyond DeepSeek-R1 and Hindi.
This work identifies a genuinely important failure mode: fact-checking models may produce correct surface verdicts in Hindi while generating zero auditable reasoning — what the authors call "hidden-unsafe certification." The distinction between veracity (P1) and auditability (P2) is conceptually clean and practically important, and the Ollama-local harness is a useful infrastructure contribution for resource-constrained deployments.
The critical problem is that every cross-lingual finding rests on N=3 paired observations. The LRDI score of 0.331, the collapse rate, the hidden-unsafe rate, and the verdict flip rate each represent exactly one event out of three — no statistical inference is possible at this scale. The paper presents these figures prominently without foregrounding that the entire cross-lingual evaluation is a three-statement pilot. Additionally, the N=150 evaluation cited throughout is English-only; the dashboard visualisations conflate this with the Hindi evaluation in a way that overstates the cross-lingual evidence. There is also an internal inconsistency: Table 3a shows all three Hindi verdicts matching English, yet the paper reports a 33.3% verdict flip rate. The LRDI keyword scoring in Hindi is unexplained — it is unclear how English logical connectives ("therefore," "however") are detected in Hindi text. The most actionable fix is running the paired evaluation on the full 150 statements, which the reported throughput suggests is feasible on local hardware.
Read full reviewShow less
This paper has a really unique main idea, and the way it frames things is the most novel and strongest part of the entire set. Identifying veracity (whether the decision is correct) from auditability (the user gets substantial reasoning in their native language), and calling the situation where the surface level decision made by the system is correct in terms of what was asked in English but there is no reasoning provided to support the decision in Hindi, "Hidden Unsafe Certification", is a great example of how to frame a previously unaddressed failure point clearly. Using the grocery shelves example, where English generates 847 characters of reasoning about the question, while Hindi answers "I cannot evaluate this" and still gives the same decision as English shows the problem with using solely the decision of the system as a benchmark. There are also several good ideas to build upon the contributions of this paper.
The first area that needs improvement is scaling up the paired comparison. Right now we have only 3 paired statements for each headline cross-lingual number (LRDI .331 and all three of the 33.3%). Each rate we see is based on only one out of the three, therefore we do not trust the rates. We need many more examples of paired comparisons to validate these numbers, at least dozens and preferably all 150 paired. While the phenomenon exists and is valid, we do not have enough evidence to make any concrete claims with our current data set.
The second area that needs some clarity is whether the authors mean to imply they used N=3 or N=150 when calculating their rates. In other words, I find it confusing that we report that there were 150 runs in both tables and then go on to use numbers generated from only three paired statements. We should label every single figure with the actual amount of statements used when generating said metrics.
Thirdly, we need to separate the potential for translation errors vs. the failure of reasoning for the model. Since Hindi is translated using machines, a lack of chain-of-thought in Hindi may be due to issues with either translation or prompting language rather than an attribute of the model. Therefore, creating a series of native speaker translations, which you list as future research along with checking the quality of the translations, will help us better understand whether or not there was a true failure in reasoning by the model.
Fourthly, we should include baselines for accuracy. English answer accuracy of 12.9% for a four-way verdict task is extremely poor. Adding a majority class baseline and noting the difficulty of the task will allow readers to put this into perspective.
Lastly, we need to trim down for signal. The paper is too long and repeats the same findings throughout multiple dashboard screens and sections related to roadmaps. If we focus on the threat model, define our metrics, present the case study, and provide overall results, this will give us a greater chance of making this key concept more impactful.
Read full reviewShow less
Cite this project
@misc{priyanka2026linguistic,
title = {{Linguistic Reasoning Drift Index (LRDI): Auditing Multilingual Misinformation Safety for the Global South}},
author = {Priyanka and Anushka and Shweta singh and Kirti saini},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/linguistic-reasoning-drift-index-lrdi-auditing-multilingual-misinformation-safety-for-the-global-south-6zdy}},
url = {https://apartresearch.com/sprints/projects/linguistic-reasoning-drift-index-lrdi-auditing-multilingual-misinformation-safety-for-the-global-south-6zdy}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …