The Veneer of safety: the fragility of India-specific harm refusal in open LLMs

Vihaan B., Ketki Banafar

Mechanistic-interpretability study showing open 7-8B LLMs internally represent India-specific harms but refuse them only superficially: canonical-anchored safeguards miss them, and a benign LoRA finetune strips that refusal on Llama-2.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

Strong submission. The question you ask is a good one and it matters: do these small open models actually understand India-specific harms like dowry coercion, sex-selective abortion, and faith healing, or do they just refuse on the surface? Splitting recognition from refusal is the right framing, and the Airavata comparison is the best part of the paper. Since it shares the Llama-2 base but has no safety tuning, it isolates the effect of safety tuning — same starting model, opposite refusal behavior. That is a clean piece of evidence.

A few things to work on, in order of importance:

1. The biggest issue is that the harm direction is fit and tested on the same small set of prompts. With about 15 prompts per category in a high-dimensional space, this inflates the Cohen's d values — the d's of 6 to 10 for Airavata are too high to be real effect sizes. The simple fix is a train/test split: fit the direction on one half of the prompts and report d on the other half. This also affects your headline coverage gap, since own_d is measured in-sample while canon_d is out-of-sample for India categories, so part of the low coverage ratio could come from that gap, not from canonical harm being a different concept. To your credit, Airavata's high coverage ratios under the same method show the gap is not only a fitting artifact, so the result still holds something real. But a split would make the claim much stronger.

2. The main claim — that recognition survives while refusal erodes — is inferred, not measured. You never re-measure the harm direction after the finetune. You point this out yourself, and it is also the cheapest experiment that would most strengthen the paper. I would run it.

3. Smaller points: the substring refusal classifier is coarse, and with 15 prompts per category the per-category rates are rough. The within-India ordering also does not match the coverage ratio (scams_fraud loads highest on the canonical axis but drops the most), which cuts against the clean story. You report all of this honestly, which is good.

On presentation: the metrics are well defined and the limitations section is unusually honest. A few small fixes would help — Figure 2 is referenced before Figure 1, the abstract is dense, and one table is split across pages.

Overall this is careful, self-aware work on a problem the field overlooks. The diagnosis is solid; the next step is to close the loop with a train/test split and a post-finetune re-measurement so the "veneer" claim is shown directly rather than inferred.

Clear document. Intro and discussion well organized. Main body has lots of graphs that are self explained but not fully discussed in the main body . Methodology is strong and research question is not only replication.

This is a strong and timely investigation of multilingual refusal mechanisms in small open language models. The project addresses an important AI-safety problem: safety behavior that appears robust in English may not transfer reliably to lower-resource languages. The use of topic-matched harmful/benign prompt pairs, three model families, three languages, external behavioral judging, confidence intervals, and causal activation intervention makes the study substantially stronger than a purely correlational representation analysis. The multi-model replication of cross-lingual refusal reduction after subtracting an English-derived direction is the most valuable contribution. The comparison between dense refusal directions and sparse-autoencoder features is also useful, particularly because the authors include a within-language stability baseline rather than over-interpreting raw cross-language feature overlap.

Several issues should be addressed before making the central claims more definitive. First, some headline statements appear stronger than the reported results. The abstract states cosine similarities of approximately 0.85–0.97 and transfer AUROCs of 0.92–0.98, while Table 1 reports a Qwen English–Hindi cosine of 0.53 and transfer AUROC of 0.76. Similarly, describing refusal as “collapsing” in every language overstates the Qwen Hindi result, where refusal changes from 0.33 to 0.17 and the available headroom is limited. These inconsistencies should be corrected, and weaker cases should be described explicitly as partial or uncertain evidence.

Second, the causal analysis would benefit from stronger controls and statistical testing. Wilson intervals on post-intervention refusal rates do not directly establish uncertainty in the paired refusal-rate difference. The authors should report paired bootstrap confidence intervals or an appropriate paired significance test for each baseline-versus-intervention comparison. Norm-matched random directions, unrelated activation directions, token-position sweeps, and systematic coherence or output-quality measurements would help demonstrate that the effect is specific to refusal rather than a broader degradation of generation behavior.

Third, the external judge is preferable to keyword matching or self-judging, but it is not human-validated. A blinded multilingual human-labelled subset, inter-rater agreement, and judge-versus-human error analysis would materially strengthen the behavioral conclusions. Completing the Vietnamese native-speaker audit is also important because translation artifacts could influence both representation geometry and refusal behavior.

Finally, the SAE result should remain framed as inconclusive rather than evidence against cross-lingual feature drift. It is based on one model, one selected layer, one top-k overlap metric, and a relatively unstable within-language baseline. Testing multiple layers, SAE widths, feature-selection thresholds, seeds, and additional model families would make this comparison much more informative.

Overall, the project is technically competent, clearly presented, and potentially valuable to multilingual AI-safety research. Correcting the numerical overstatements and adding stronger causal, human-evaluation, and robustness controls would significantly increase confidence in the conclusions.

scores

| Criterion | Score | Rationale |

| ----------------------------- | ------: | ------------------------------------------------------------------------------------------------------------------------------------ |

| Impact Potential & Innovation | 4.0 | Important multilingual safety question with a useful multi-model causal contribution |

| Execution Quality | 3.5 | Strong hackathon execution, but judge validation, intervention controls, paired statistics, and translation audits remain incomplete |

| Presentation & Clarity | 4.0 | Clear and concise overall, though the abstract and figure language overstate some weaker Table 1 results |

Cite this work

@misc {

title={

(HckPrj) The Veneer of safety: the fragility of India-specific harm refusal in open LLMs

},

author={

Vihaan B., Ketki Banafar

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.