Skip to content
Sprint projectJun 21, 2026Chittagong, Bangladesh

StyleSwitch-BN: Auditing Bengali LLM Safety Across Real-World Writing Styles

Naymul Islam · Team StyleSwitch-BN

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: StyleSwitch-BN: Auditing Bengali LLM Safety Across Real-World Writing Styles

Code (opens in new tab)
Share

Most non-English LLM safety tests translate one English harmful prompt and check for a refusal, assuming a language has a single voice. Bengali does not: people write it as polished news prose, casual Banglish chat, or stiff official language. We asked whether keeping the harmful request identical and changing only the writing style changes the model's safety response. It does. Across six frontier models, the same request gets an unsafe answer 61.0% of the time in news-style Bengali but only 44.6% in casual code-mixed Bengali, a 16-point gap from style alone with no jailbreak tricks. We release labels-only data, reproduction code, figures, and a dashboard.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The framing here is genuinely novel and practically important. Most multilingual safety work asks "does safety degrade across languages?" This paper asks "does safety degrade across everyday writing styles within a language?" — and finds a 16.4-point gap between news-style and casual Bengali, with formal writing being riskier. The finding that the English baseline (58.9%) sits close to news-style Bengali (61.0%), while casual Bengali (44.6%) is actually safer, upends the usual narrative. The diglossia framing (Section 6) correctly situates this as a structural linguistic phenomenon, not an adversarial trick.

    Several suggestions for strengthening: (1) The bootstrapped confidence intervals are appropriate, but with 1,530 labeled responses across six models the effective sample per model-style cell may be small — reporting per-model cell sizes would help readers assess how much variation comes from individual models. (2) The "Style Robustness Risk Score" composite metric (0.5xUER + 0.3xgap + 0.2xDCR) is somewhat arbitrary weighting — a sensitivity analysis showing the ranking is stable across different weight choices would make the leaderboard more trustworthy. (3) Single annotator is the main threat to validity — even a small inter-rater reliability check on a subset would substantially strengthen the results.

    The responsible-release approach (labels only, no prompts, no model answers) is commendable and sets a good precedent for others working in this space.

    Read full reviewShow less
  2. This is an important AI safety project and should be taken into consideration when models are deployed in different markets, languages, and cultural contexts. The authors provide evidence across multiple models and writing styles that safety behavior can vary depending on prompt style, making this a relevant multilingual AI safety concern. The finding that news-style Bengali prompts appear to weaken safety guardrails is especially interesting. However, the dataset could have been larger and more diverse, and there is limited evidence of external validation through independent reviewers. Overall, this is a promising safety audit, but the evidence is not conclusive.

Cite this project

@misc{islam2026styleswitchbn,
  title = {{StyleSwitch-BN: Auditing Bengali LLM Safety Across Real-World Writing Styles}},
  author = {Naymul Islam},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/styleswitchbn-auditing-bengali-llm-safety-across-realworld-writing-styles-6f0a}},
  url = {https://apartresearch.com/sprints/projects/styleswitchbn-auditing-bengali-llm-safety-across-realworld-writing-styles-6f0a}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026