One Direction, Many Languages: Causal Cross-Lingual Refusal Transfer Across Small Open Models
Aradhya Goel, Bhoomika Gupta
This work tests whether the internal "refusal direction" that small language models use to reject harmful prompts is the same across English, Hindi, and Vietnamese. Using three open models (Gemma-2-2B, Llama-3.2-1B, and Qwen2.5-1.5B) and topic-matched prompt pairs in each language, we show that the direction transfers across languages both geometrically and causally: subtracting the English direction reduces refusal not only in English but also in Hindi and Vietnamese, and this effect replicates across all three models. In contrast, the sparse autoencoder features do not show reliable cross-lingual transfer, indicating they are a noisier probe than the dense direction. Overall, the work demonstrates that cross-lingual refusal in these models is a shared mechanism rather than a set of independent, language-specific ones.
This is methodologically careful work. The multi-model causal replication is the paper's main contribution: showing that ablating the English refusal direction reduces refusal in Hindi and Vietnamese across three model families (Gemma, Llama, Qwen) is stronger evidence than any single-model study provides. The measurement-failure section is also genuinely useful — demonstrating that keyword classifiers agree with an external judge at chance (κ≈0) and that self-judging mislabels Vietnamese outputs systematically is a direct warning to the broader field.
Clarify and cite the SAE drift claim being tested. The work frames its SAE result as a negative finding against a "features drift" narrative, but cites only Lieberum et al. (Gemma Scope) for the SAE tool itself — not for any cross-lingual drift claim. The negative result becomes substantially more impactful if the paper either cites the specific source making the drift claim, or reframes the SAE section as establishing a within-language stability baseline rather than rebutting an established claim.
Extend the SAE analysis to all three models. The causal direction result replicates across all three model families but the SAE Jaccard analysis is Gemma-only. Given that the null result is the paper's most novel finding, showing it holds across Llama and Qwen too would make it considerably more convincing.
Its' a strong safety relevance project. The idea that different languages may share a common refusal failure mode is interesting and important. The hackathon scope is good, and the authors are transparent about using smaller models. The main limitation is evidence strength: it would be useful to compare against larger model classes, validate across a larger multilingual dataset, and include more human-evaluated labels to confirm the conclusion is robust beyond small models and automated judging.
Great work. Project extends existing research and proves hypothesis. Methodology is strong and supports conclusions. Writing is easy to follow and well structured.
This is a strong and timely investigation of multilingual refusal mechanisms in small open language models. The project addresses an important AI-safety problem: safety behavior that appears robust in English may not transfer reliably to lower-resource languages. The use of topic-matched harmful/benign prompt pairs, three model families, three languages, external behavioral judging, confidence intervals, and causal activation intervention makes the study substantially stronger than a purely correlational representation analysis. The multi-model replication of cross-lingual refusal reduction after subtracting an English-derived direction is the most valuable contribution. The comparison between dense refusal directions and sparse-autoencoder features is also useful, particularly because the authors include a within-language stability baseline rather than over-interpreting raw cross-language feature overlap.
Several issues should be addressed before making the central claims more definitive. First, some headline statements appear stronger than the reported results. The abstract states cosine similarities of approximately 0.85–0.97 and transfer AUROCs of 0.92–0.98, while Table 1 reports a Qwen English–Hindi cosine of 0.53 and transfer AUROC of 0.76. Similarly, describing refusal as “collapsing” in every language overstates the Qwen Hindi result, where refusal changes from 0.33 to 0.17 and the available headroom is limited. These inconsistencies should be corrected, and weaker cases should be described explicitly as partial or uncertain evidence.
Second, the causal analysis would benefit from stronger controls and statistical testing. Wilson intervals on post-intervention refusal rates do not directly establish uncertainty in the paired refusal-rate difference. The authors should report paired bootstrap confidence intervals or an appropriate paired significance test for each baseline-versus-intervention comparison. Norm-matched random directions, unrelated activation directions, token-position sweeps, and systematic coherence or output-quality measurements would help demonstrate that the effect is specific to refusal rather than a broader degradation of generation behavior.
Third, the external judge is preferable to keyword matching or self-judging, but it is not human-validated. A blinded multilingual human-labelled subset, inter-rater agreement, and judge-versus-human error analysis would materially strengthen the behavioral conclusions. Completing the Vietnamese native-speaker audit is also important because translation artifacts could influence both representation geometry and refusal behavior.
Finally, the SAE result should remain framed as inconclusive rather than evidence against cross-lingual feature drift. It is based on one model, one selected layer, one top-k overlap metric, and a relatively unstable within-language baseline. Testing multiple layers, SAE widths, feature-selection thresholds, seeds, and additional model families would make this comparison much more informative.
Overall, the project is technically competent, clearly presented, and potentially valuable to multilingual AI-safety research. Correcting the numerical overstatements and adding stronger causal, human-evaluation, and robustness controls would significantly increase confidence in the conclusions.
Score
| Criterion | Score | Rationale |
| Impact Potential & Innovation | 4.0 | Important multilingual safety question with a useful multi-model causal contribution |
| Execution Quality | 3.5 | Strong hackathon execution, but judge validation, intervention controls, paired statistics, and translation audits remain incomplete |
| Presentation & Clarity | 4.0 | Clear and concise overall, though the abstract and figure language overstate some weaker Table 1 results |
The most valuable part is the framing: refusal robustness in lower-resource languages like Hindi and Vietnamese is a real and under-studied deployment problem, and the paper is a useful proof of concept that the refusal direction can be probed and steered across them. The set of experiments is diverse, covering geometry, causal intervention, SAE features, and judging methodology in one pipeline. The natural next step is to verify these results at larger scale.
The causal evidence is one-sided. Subtracting the direction lowers refusal on harmful prompts, but the other half of the test is missing: adding the direction on the benign prompts, to see whether it induces refusal on harmless requests. The existing amplification control was run on harmful prompts already at ceiling, so it had no headroom and showed nothing; the benign prompts are where refusal can actually go up. Showing both arms (subtract makes the model comply on harmful prompts, add makes it refuse harmless ones) would establish that the direction genuinely controls refusal rather than just nudging the model toward compliance. It would also help to add a coherence or output-quality check on the post-subtraction generations, so it is clear the model is genuinely complying rather than degrading into broken text; the current reading-the-outputs check is qualitative and harmful-side only.
Worth verifying at larger scale. All three models are in the 1 to 2B range, where the non-English refusal baselines are already low and leave little headroom (Qwen refuses only about a third of Hindi prompts before any intervention). Running the same pipeline on a few larger models would test whether the transfer holds where refusal is firmly established in every language, and would make the cross-lingual claim considerably more robust.
Cite this work
@misc {
title={
(HckPrj) One Direction, Many Languages: Causal Cross-Lingual Refusal Transfer Across Small Open Models
},
author={
Aradhya Goel, Bhoomika Gupta
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


