Linguistic Reasoning Drift Index (LRDI): Auditing Multilingual Misinformation Safety for the Global South

Priyanka, Anushka, Shweta singh, Kirti saini

We present the Linguistic Reasoning Drift Index (LRDI), an open-source audit framework for multilingual AI safety in the Global South. LRDI evaluates open-weight reasoning models on English and Hindi misinformation prompts, detects reasoning collapse and hidden-unsafe cases, and visualizes results in a Streamlit dashboard. Our pilot with DeepSeek-R1:8b on Ollama shows 33.3% of Hindi prompts lose substantive chain-of-thought while matching English verdicts, revealing that verdict-only benchmarks can falsely certify safety. The pipeline is local, reproducible, and designed for deployment-side audits by regional institutions.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

The proposed Linguistic Reasoning Drift Index is a great attempt to quantify cross-lingual reasoning degradation, and the emphasis on auditability rather than just prediction accuracy is a valuable perspective for AI safety.

The project would be substantially strengthened by:

- Expanding the paired multilingual evaluation from 3 examples to hundreds or thousands before making deployment or policy claims.

- Validating LRDI against human judgments of reasoning quality, rather than relying primarily on heuristic proxies such as response length and logical connective density.

- Evaluating multiple models and multiple Indian languages to demonstrate that the observed reasoning drift generalizes beyond DeepSeek-R1 and Hindi.

This work identifies a genuinely important failure mode: fact-checking models may produce correct surface verdicts in Hindi while generating zero auditable reasoning — what the authors call "hidden-unsafe certification." The distinction between veracity (P1) and auditability (P2) is conceptually clean and practically important, and the Ollama-local harness is a useful infrastructure contribution for resource-constrained deployments.

The critical problem is that every cross-lingual finding rests on N=3 paired observations. The LRDI score of 0.331, the collapse rate, the hidden-unsafe rate, and the verdict flip rate each represent exactly one event out of three — no statistical inference is possible at this scale. The paper presents these figures prominently without foregrounding that the entire cross-lingual evaluation is a three-statement pilot. Additionally, the N=150 evaluation cited throughout is English-only; the dashboard visualisations conflate this with the Hindi evaluation in a way that overstates the cross-lingual evidence. There is also an internal inconsistency: Table 3a shows all three Hindi verdicts matching English, yet the paper reports a 33.3% verdict flip rate. The LRDI keyword scoring in Hindi is unexplained — it is unclear how English logical connectives ("therefore," "however") are detected in Hindi text. The most actionable fix is running the paired evaluation on the full 150 statements, which the reported throughput suggests is feasible on local hardware.

This paper has a really unique main idea, and the way it frames things is the most novel and strongest part of the entire set. Identifying veracity (whether the decision is correct) from auditability (the user gets substantial reasoning in their native language), and calling the situation where the surface level decision made by the system is correct in terms of what was asked in English but there is no reasoning provided to support the decision in Hindi, "Hidden Unsafe Certification", is a great example of how to frame a previously unaddressed failure point clearly. Using the grocery shelves example, where English generates 847 characters of reasoning about the question, while Hindi answers "I cannot evaluate this" and still gives the same decision as English shows the problem with using solely the decision of the system as a benchmark. There are also several good ideas to build upon the contributions of this paper.

The first area that needs improvement is scaling up the paired comparison. Right now we have only 3 paired statements for each headline cross-lingual number (LRDI .331 and all three of the 33.3%). Each rate we see is based on only one out of the three, therefore we do not trust the rates. We need many more examples of paired comparisons to validate these numbers, at least dozens and preferably all 150 paired. While the phenomenon exists and is valid, we do not have enough evidence to make any concrete claims with our current data set.

The second area that needs some clarity is whether the authors mean to imply they used N=3 or N=150 when calculating their rates. In other words, I find it confusing that we report that there were 150 runs in both tables and then go on to use numbers generated from only three paired statements. We should label every single figure with the actual amount of statements used when generating said metrics.

Thirdly, we need to separate the potential for translation errors vs. the failure of reasoning for the model. Since Hindi is translated using machines, a lack of chain-of-thought in Hindi may be due to issues with either translation or prompting language rather than an attribute of the model. Therefore, creating a series of native speaker translations, which you list as future research along with checking the quality of the translations, will help us better understand whether or not there was a true failure in reasoning by the model.

Fourthly, we should include baselines for accuracy. English answer accuracy of 12.9% for a four-way verdict task is extremely poor. Adding a majority class baseline and noting the difficulty of the task will allow readers to put this into perspective.

Lastly, we need to trim down for signal. The paper is too long and repeats the same findings throughout multiple dashboard screens and sections related to roadmaps. If we focus on the threat model, define our metrics, present the case study, and provide overall results, this will give us a greater chance of making this key concept more impactful.

Cite this work

@misc {

title={

(HckPrj) Linguistic Reasoning Drift Index (LRDI): Auditing Multilingual Misinformation Safety for the Global South

},

author={

Priyanka, Anushka, Shweta singh, Kirti saini

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.