Skip to content
Sprint projectJun 21, 2026Delhi, India

LatentGuard: Mitigating Multilingual Safety Bypass via Mid-Layer Latent Steering

Aishwarya Mukherjee, Surjit Chowdhary · Team AI Risk Rangers

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: LatentGuard: Mitigating Multilingual Safety Bypass via Mid-Layer Latent Steering

More on huggingface.co (opens in new tab)
Share

Modern LLMs rely heavily on Reinforcement Learning from Human Feedback (RLHF) for safety alignment, a process that remains overwhelmingly English-centric. Consequently, when prompted in low-resource Indic languages like Bengali, models suffer from Safety Drift, systematically failing to refuse harmful inputs or exhibiting high refusal variance. Conventional mitigations attempt cross-lingual fine-tuning using synthetically translated data. However, this approach triggers Knowledge Collapse, while superficial linguistic fluency remains, factual integrity and geometric safety boundaries dissolve because the underlying representations are corrupted by poor token fragmentation. Recursive fine-tuning on these shattered representations accelerates degradation rather than preventing it. Using mechanistic interpretability on Gemma 2 2B Instruct, we quantify this internal geometric decay. Our analysis reveals a late-stage safety bypass: the model successfully maps Bengali malicious intent to English safety anchors in its mid-layers, but experiences severe cross-lingual drift immediately prior to output generation. To address this structural failure, we introduce LatentGuard, a training-free inference intervention. By applying mid-layer latent steering, we actively project drifting Bengali representations back into the stable English safety subspace right before divergence occurs. In our evaluations, LatentGuard reduced the Bengali Attack Success Rate (ASR) on safety benchmarks by 20%, while maintaining fluency. This approach dynamically restores cross-lingual safety boundaries without the computational costs, data scarcity, or Knowledge Collapse risks inherent to synthetic fine-tuning.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. LatentGuard addresses a real and underappreciated failure mode: safety alignment in LLMs degrades systematically for low-resource Indic languages like Bengali, and synthetic fine-tuning — the current standard mitigation — risks Knowledge Collapse by forcing models to memorize broken token representations. The paper's clearest contribution is the "illusion of alignment" finding: Gemma-2-2B-IT successfully maps Bengali harmful intent to English safety anchors at Layer 13 (cross-lingual drift = 0.047, ~95% cosine similarity), but this alignment unravels catastrophically by Layer 25 (drift = 0.510, ~49% similarity). This non-monotonic trajectory reframes the problem from "the model doesn't understand Bengali" to "the model understands but structurally diverges before output generation" — a meaningful mechanistic reframing with practical implications. The training-free inference intervention follows naturally from this diagnostic, and the three-phase structure (behavioral profiling → latent diagnostics → steering intervention) is well-organized and easy to follow.

    The primary limitation is scope. The failure cohort is 35 prompts from a single model and a single language. Whether the Late-Stage Safety Bypass generalizes to Hindi, Tamil, Telugu, or Marathi — or to Llama-3 or Mistral — is entirely open. The headline claim ("multilingual safety boundaries") outpaces what 35 Bengali prompts can support. Additionally, refusal detection relies entirely on pattern-matched indicator strings rather than human verification or LLM judging, which means the reported 20pp improvement (60% → 80% refusal rate) may partially reflect the model emitting English-language boilerplate rather than genuine cross-lingual safety restoration. A stratified human evaluation of even 10–15 steered responses would substantially strengthen the result.

    The Language Bleed finding — steered responses arriving in English despite Bengali input — is disclosed honestly and is the paper's most practically significant limitation. The orthogonal projection approach suggested in future work (projecting only the safety-relevant component of the steering vector, orthogonal to the English-Bengali language direction) is exactly the right next step and should be the first priority for follow-on work.

    Read full reviewShow less
  2. I really love this project. This is ambitious and impressively complete - the only project to run the full arc of diagnose, localize, intervene, and ablate, and you do it with real statistics (t(34)=4.2, p<0.001) and a clean layer x alpha ablation. The mechanistic story is well supported (drift bottoming out at layer 13, then diverging sharply by layer 25), and the "Language Bleed" observation is excellent, mature self-critique: noticing the steered model refuses in English rather than Bengali, recognizing this breaks accessibility for the very users it's meant to protect, and proposing orthogonal projection as the fix. The info-hazard notice and open dataset are good practice.

    Strong, creative work - the orthogonal-projection version that disentangles safety from language is the most valuable next step.

  3. This paper admits that multilingual safety disparities and latent steering are established in the research, and this is a case study of a specific English/Bengali gap in a single gemma model that was studied.

    For each of the 150 parallel prompt pairs the metrics used were refusal mismatch, token fragmentation and confident hallucination. Only the first of these are actually a measure of safety the other two are more about low capability. So to improve this work I would be interested to see the work done with a variety of models, with more diverse prompts.

    The current steering vector is computed from the same 35-prompt failure cohort used for evaluation, which risks overfitting.

    In methods it is said "To create our target distribution for Phase 2 and 3, we analyzed a 150-prompt baseline

    and isolated a 35-prompt failure cohort. instances of verified asymmetric failure where English

    prompts triggered refusal while the Bengali translation triggered compliance."

    Then later in the steering results, the paper says the intervention was applied to the 35 failure prompts, and the baseline before steering refusal rate was already 60.0%. This was confusing because I would have espected the baseline to be 0% refused, 100% harmful compliance as per the cohort definition. So it is important that the paper maintains terms like failure cohort consistently and seperate out output degeneration, refusal, steering, token fragmentation so a cohort flow diagram would really help to see what happened, maybe like a sankey diagram.

    I would be excited to see validation experiments done, these negative controls, especially a random vector, a shuffled prompt-label vector, a same-norm vector unrelated to Bengali-English differences, would be valuable in making this result strong.

    Read full reviewShow less

Cite this project

@misc{mukherjee2026latentguard,
  title = {{LatentGuard: Mitigating Multilingual Safety Bypass via Mid-Layer Latent Steering}},
  author = {Aishwarya Mukherjee and Surjit Chowdhary},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/latentguard-mitigating-multilingual-safety-bypass-via-midlayer-latent-steering-zpoz}},
  url = {https://apartresearch.com/sprints/projects/latentguard-mitigating-multilingual-safety-bypass-via-midlayer-latent-steering-zpoz}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026