Skip to content
Sprint projectJun 22, 2026Pakistan

Safeswitch: A Localized Benchmark for LLM Safety in Urdu and Pashto Across Language Forms and Prompting Strategies

Muhammad Ali · Team SafeSwitch

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Safeswitch: A Localized Benchmark for LLM Safety in Urdu and Pashto Across Language Forms and Prompting Strategies

Presentation

Presentation: Safeswitch: A Localized Benchmark for LLM Safety in Urdu and Pashto Across Language Forms and Prompting Strategies

Code (opens in new tab)
Share

SafeSwitch is a small benchmark that tests whether LLMs stay safe when harmful requests are written in Urdu, Pashto, romanized script, or code-switching instead of English. It uses 20 seed prompts across four safety categories (medical, scam, hate, cyber), rewritten by human into seven language forms, and three prompting strategies (zero-shot, chain-of-thought, persona). Three models were tested: GPT-5.4, Qwen3.7-Plus, and Gemma-3-4B. Responses were scored with a two-signal pipeline: rule-based refusal detection plus a GPT-5.5 judge that labels both harm and prompt comprehension. Across 588 responses, harmful output concentrated in non-English forms (16 of 17 under zero-shot), and the SLM's low harm rate on Pashto turned out to be comprehension failure, not safety. The result is a reproducible dataset, scoring code, and metrics for Urdu and Pashto LLM safety.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Here is some constructive feedback to help strengthen the project, categorized by methodology, scaling, and presentation:

    1. Strengthening the Methodology

    - Validate the LLM Judge: You rightly noted in the limitations that the GPT-5.5 judge lacks human validation. Since frontier models can also struggle with Pashto and code-switched nuances, validating a subset of the judge's verdicts with native speakers is critical. Even a small sample (e.g., 50-100 responses) evaluated by humans would give your judge's metrics much more credibility.

    - Translation Quality Assurance: The benchmark currently relies on unvalidated translations. Before expanding the dataset too much, it would be highly beneficial to have native speakers (especially for Pashto and the romanized/code-switched forms) verify that the seeds retain their intended harmful semantics and sound natural.

    2. Scaling for Statistical Significance

    - Expand the Prompting Strategy Subset: Your findings on Chain-of-Thought (CoT) and Persona prompting are fascinating—specifically that they push harm entirely into non-English forms for the small model. However, because this was tested on only 4 seeds (yielding very small sample sizes), the results are currently exploratory. Running the full 20-seed set through the CoT and Persona strategies should be a priority for the next iteration to see if those trends hold with statistical significance.

    - Increase the Base Seed Count: 20 seeds across four categories is a great proof-of-concept for a hackathon, but scaling this to 50–100 seeds per category would make the benchmark robust enough for a major conference publication.

    3. Presentation and Readability

    - Include Qualitative Examples: The paper would benefit greatly from a few real examples of the model outputs. Showing exactly how a model safely deflected in English but fell for the exact same Scam prompt in Roman Urdu or Pashto would make your findings much more concrete for the reader.

    - Clarify the Refusal vs. Harm Relationship: You mention that Refusal Rate (RR) and Harmful Response Rate (HRR) are not complementary. A brief example illustrating a "safe deflection" (not harmful, but missing refusal keywords) versus a "partial refusal" (contains refusal keywords but still gives harmful info) would clarify this distinction beautifully.

    Overall, it's a highly relevant and well-designed benchmark for a hackathon timeline.

    Read full reviewShow less
  2. Great to see the somewhat novel approach of combining personas with chain-of-thought. And tracking of prompt comprehension, an easily overlooked aspect

    While the novelty and depth of the paper are great, using only 4 seed prompts limits the solidity of the outcomes and conclusions. Statistically significant outcomes can be trusted more

    Very intuitive way of presenting the data, organizing the paper around problem statement. Good callout and delineation between actual safety and comprehension failure

  3. The most important next step is scaling the seed set, because several of the paper's most interesting findings, particularly the prompting-by-language interaction and the per-form harm rates, currently rest on so few data points that they read as pilot signals rather than established results. Overall this is a focused and reproducible piece of work that identifies a real gap and opens a well-defined research direction, and the three future work threads (prompting techniques, language coverage, alignment) are exactly the right priorities.

  4. This is a clear and regionally well-motivated AI safety submission. SafeSwitch targets an important evaluation gap: LLM safety behavior in Urdu and Pashto, including native-script, romanized, and code-switched forms that reflect how users in Pakistan, Afghanistan, and diaspora communities actually write. The benchmark design is clean, crossing seven language forms with four safety categories and three prompting strategies, and the inclusion of a comprehension signal is especially valuable because it prevents incomprehension from being mistaken for safe refusal. The finding that harmful responses cluster in non-English forms, especially for scam/fraud prompts, is safety-relevant and worth investigating further.

    That said, the broader area of multilingual safety gaps, code-switching attacks, language-specific safety benchmarks, and refusal robustness is already well studied in the research literature, including work such as "All Languages Matter: On the Multilingual Safety of LLMs", "Multilingual Jailbreak Challenges in Large Language Models", and recent code-switching red-teaming studies. The main contribution here is therefore localized coverage and careful benchmark construction rather than a wholly novel method. The main limitations are scale and validation: the benchmark uses only 20 seed prompts, the chain-of-thought and persona analyses use only 4 seeds, several effects rest on very small counts, and harm/comprehension labels come from a single GPT-5.5 judge without completed human spot-checking. Pashto prompt quality and romanization would also benefit from independent native-speaker review.

    Overall, this is a solid, useful hackathon benchmark with a clear path forward: expand the seed set, add native-speaker and human-judge validation, sample multiple generations, include more models, and then re-test the prompting-by-language interaction at scale. I encourage the author to continue developing this work, especially because Urdu, Pashto, Roman Pashto, and code-switched safety evaluation remain important and underserved areas.

    Read full reviewShow less

Cite this project

@misc{ali2026safeswitch,
  title = {{Safeswitch: A Localized Benchmark for LLM Safety in Urdu and Pashto Across Language Forms and Prompting Strategies}},
  author = {Muhammad Ali},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/safeswitch-a-localized-benchmark-for-llm-safety-in-urdu-and-pashto-across-language-forms-and-prompting-strategies-w9yl}},
  url = {https://apartresearch.com/sprints/projects/safeswitch-a-localized-benchmark-for-llm-safety-in-urdu-and-pashto-across-language-forms-and-prompting-strategies-w9yl}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026