One Direction, Many Languages: Causal Cross-Lingual Refusal Transfer Across Small Open Models

Aradhya Goel, Bhoomika Gupta

This work tests whether the internal "refusal direction" that small language models use to reject harmful prompts is the same across English, Hindi, and Vietnamese. Using three open models (Gemma-2-2B, Llama-3.2-1B, and Qwen2.5-1.5B) and topic-matched prompt pairs in each language, we show that the direction transfers across languages both geometrically and causally: subtracting the English direction reduces refusal not only in English but also in Hindi and Vietnamese, and this effect replicates across all three models. In contrast, the sparse autoencoder features do not show reliable cross-lingual transfer, indicating they are a noisier probe than the dense direction. Overall, the work demonstrates that cross-lingual refusal in these models is a shared mechanism rather than a set of independent, language-specific ones.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is methodologically careful work. The multi-model causal replication is the paper's main contribution: showing that ablating the English refusal direction reduces refusal in Hindi and Vietnamese across three model families (Gemma, Llama, Qwen) is stronger evidence than any single-model study provides. The measurement-failure section is also genuinely useful — demonstrating that keyword classifiers agree with an external judge at chance (κ≈0) and that self-judging mislabels Vietnamese outputs systematically is a direct warning to the broader field.

Clarify and cite the SAE drift claim being tested. The work frames its SAE result as a negative finding against a "features drift" narrative, but cites only Lieberum et al. (Gemma Scope) for the SAE tool itself — not for any cross-lingual drift claim. The negative result becomes substantially more impactful if the paper either cites the specific source making the drift claim, or reframes the SAE section as establishing a within-language stability baseline rather than rebutting an established claim.

Extend the SAE analysis to all three models. The causal direction result replicates across all three model families but the SAE Jaccard analysis is Gemma-only. Given that the null result is the paper's most novel finding, showing it holds across Llama and Qwen too would make it considerably more convincing.

Its' a strong safety relevance project. The idea that different languages may share a common refusal failure mode is interesting and important. The hackathon scope is good, and the authors are transparent about using smaller models. The main limitation is evidence strength: it would be useful to compare against larger model classes, validate across a larger multilingual dataset, and include more human-evaluated labels to confirm the conclusion is robust beyond small models and automated judging.

Great work. Project extends existing research and proves hypothesis. Methodology is strong and supports conclusions. Writing is easy to follow and well structured.

This is a strong and timely investigation of multilingual refusal mechanisms in small open language models. The project addresses an important AI-safety problem: safety behavior that appears robust in English may not transfer reliably to lower-resource languages. The use of topic-matched harmful/benign prompt pairs, three model families, three languages, external behavioral judging, confidence intervals, and causal activation intervention makes the study substantially stronger than a purely correlational representation analysis. The multi-model replication of cross-lingual refusal reduction after subtracting an English-derived direction is the most valuable contribution. The comparison between dense refusal directions and sparse-autoencoder features is also useful, particularly because the authors include a within-language stability baseline rather than over-interpreting raw cross-language feature overlap.

Several issues should be addressed before making the central claims more definitive. First, some headline statements appear stronger than the reported results. The abstract states cosine similarities of approximately 0.85–0.97 and transfer AUROCs of 0.92–0.98, while Table 1 reports a Qwen English–Hindi cosine of 0.53 and transfer AUROC of 0.76. Similarly, describing refusal as “collapsing” in every language overstates the Qwen Hindi result, where refusal changes from 0.33 to 0.17 and the available headroom is limited. These inconsistencies should be corrected, and weaker cases should be described explicitly as partial or uncertain evidence.

Second, the causal analysis would benefit from stronger controls and statistical testing. Wilson intervals on post-intervention refusal rates do not directly establish uncertainty in the paired refusal-rate difference. The authors should report paired bootstrap confidence intervals or an appropriate paired significance test for each baseline-versus-intervention comparison. Norm-matched random directions, unrelated activation directions, token-position sweeps, and systematic coherence or output-quality measurements would help demonstrate that the effect is specific to refusal rather than a broader degradation of generation behavior.

Third, the external judge is preferable to keyword matching or self-judging, but it is not human-validated. A blinded multilingual human-labelled subset, inter-rater agreement, and judge-versus-human error analysis would materially strengthen the behavioral conclusions. Completing the Vietnamese native-speaker audit is also important because translation artifacts could influence both representation geometry and refusal behavior.

Finally, the SAE result should remain framed as inconclusive rather than evidence against cross-lingual feature drift. It is based on one model, one selected layer, one top-k overlap metric, and a relatively unstable within-language baseline. Testing multiple layers, SAE widths, feature-selection thresholds, seeds, and additional model families would make this comparison much more informative.

Overall, the project is technically competent, clearly presented, and potentially valuable to multilingual AI-safety research. Correcting the numerical overstatements and adding stronger causal, human-evaluation, and robustness controls would significantly increase confidence in the conclusions.

Score

| Criterion | Score | Rationale |

| Impact Potential & Innovation | 4.0 | Important multilingual safety question with a useful multi-model causal contribution |

| Execution Quality | 3.5 | Strong hackathon execution, but judge validation, intervention controls, paired statistics, and translation audits remain incomplete |

| Presentation & Clarity | 4.0 | Clear and concise overall, though the abstract and figure language overstate some weaker Table 1 results |

The most valuable part is the framing: refusal robustness in lower-resource languages like Hindi and Vietnamese is a real and under-studied deployment problem, and the paper is a useful proof of concept that the refusal direction can be probed and steered across them. The set of experiments is diverse, covering geometry, causal intervention, SAE features, and judging methodology in one pipeline. The natural next step is to verify these results at larger scale.

The causal evidence is one-sided. Subtracting the direction lowers refusal on harmful prompts, but the other half of the test is missing: adding the direction on the benign prompts, to see whether it induces refusal on harmless requests. The existing amplification control was run on harmful prompts already at ceiling, so it had no headroom and showed nothing; the benign prompts are where refusal can actually go up. Showing both arms (subtract makes the model comply on harmful prompts, add makes it refuse harmless ones) would establish that the direction genuinely controls refusal rather than just nudging the model toward compliance. It would also help to add a coherence or output-quality check on the post-subtraction generations, so it is clear the model is genuinely complying rather than degrading into broken text; the current reading-the-outputs check is qualitative and harmful-side only.

Worth verifying at larger scale. All three models are in the 1 to 2B range, where the non-English refusal baselines are already low and leave little headroom (Qwen refuses only about a third of Hindi prompts before any intervention). Running the same pipeline on a few larger models would test whether the transfer holds where refusal is firmly established in every language, and would make the cross-lingual claim considerably more robust.

Cite this work

@misc {

title={

(HckPrj) One Direction, Many Languages: Causal Cross-Lingual Refusal Transfer Across Small Open Models

},

author={

Aradhya Goel, Bhoomika Gupta

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.