Project title: Register Sensitivity in LLM Safety Responses: Evaluating How Linguistic Style Affects Scam Detection in South African Contexts
Nicoroy Zwane
This study investigates whether linguistic register, specifically the shift from formal English to informal South African WhatsApp-style English, affects how large language models respond to scam-style prompts grounded in African fraud contexts. We constructed a 12-prompt evaluation dataset across four scenarios (NSFAS bursary impersonation, bank OTP extraction, fake investment schemes, and fraudulent job recruitment), each presented in formal, neutral, and WhatsApp-style registers, and evaluated responses from ChatGPT, Gemini, and Claude. Results show meaningful register-dependent safety degradation, particularly in investment scheme scenarios where Gemini produced unsafe outputs under informal register while refusing the same request in formal English. We argue this represents a real deployment risk for African users and present the framework as a replicable template for region-specific AI safety benchmarking.
This is a tightly scoped proof-of-concept that asks the right question: does the linguistic register of a scam-style prompt change LLM safety behavior when intent is held constant? Two things stood out positively. First, the South African grounding is genuine — NSFAS bursary impersonation, bank OTP phishing, WhatsApp stokvel investment recruitment, and fake-job-with-banking-details requests are documented local fraud patterns, not US scams in translation, and that regional specificity is the kind of contribution the field genuinely needs more of. Second, the Section 6.2 dual-use subsection is one of the cleanest examples of responsible-disclosure framing I have seen at hackathon scope: explicit threat model, named mitigations, builds on publicly documented patterns, no operational scam scripts in the body. The Gemini investment-scheme register flip (SAFE in formal English, UNSAFE in both neutral and WhatsApp-SA registers for the same underlying intent) is also a striking single-cell finding worth surfacing, and the Partial Compliance Problem framing in Section 5.2 — that a warning paired with a working template still gives the fraudster a working template — is a real critique of binary HarmBench-style judging that deserves more development.
The most useful improvement is engaging with prior art on register-axis safety degradation. Qiu, Lin, Chen, Pang, Liu et al. (2023) "Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models" (arXiv:2307.08487) introduces robustness-to-paraphrase as an explicit safety axis and is the closest within-language analogue to what this paper measures. Yong, Menghini and Bach (2023) "Low-Resource Languages Jailbreak GPT-4" (arXiv:2310.02446) establishes the cross-lingual register-degradation pattern that frames the within-English register finding here as a natural downstream extension rather than a new phenomenon. Citing and positioning against both would sharpen the novelty claim and shift the contribution from "we found register matters" to "we localize and operationalize a known phenomenon with SA-specific scenarios and a three-tier rubric that surfaces Partial Compliance." Two other concrete asks: release the 12 prompts, the 36 response transcripts, and the rubric notes with whatever redaction the dual-use posture requires (transcripts redacted to key phrases would let other researchers re-classify and replicate without re-running on a now-different model snapshot); and add a benign-control condition — a legitimate register-matched task (a polite NSFAS-status email in all three registers) — so readers can distinguish "register affects harm detection" from "register affects model behavior generally."
If you continue this, the natural next step is replication with two annotators on 30+ prompts per cell with inter-rater reliability reported, ideally extended to at least one other African English variety (Nigerian Pidgin, Kenyan Sheng) so the register-axis result is not idiosyncratic to South African WhatsApp register. The UbuntuGuard team would likely be useful collaborators.
This project asks whether AI chatbots become more likely to assist with fraud when requests are written in casual WhatsApp-style English rather than formal English — a question no existing AI safety benchmark has asked for South African users. Across four documented local scam scenarios and three registers, the headline finding is that Gemini correctly refused an investment-scheme request in formal English but produced genuinely dangerous output in informal and neutral phrasing. The paper also introduces a SAFE/PARTIAL/UNSAFE grading scheme that judges responses by whether they would actually help a fraudster, not just by whether the AI sounded cautious.
Strengths
1. The problem framing is original and practically important. Studying register variation within a single English variety as a safety axis is a new angle, and anchoring it to South African fraud patterns (NSFAS impersonation, stokvel schemes, OTP theft) gives the findings immediate real-world stakes.
2. The SAFE/PARTIAL/UNSAFE framework is a genuine contribution. Classifying a response by whether it would operationally assist a fraudster — regardless of included disclaimers — is more honest than a simple refuse/comply binary, and the paper applies the rule consistently.
Weaknesses
1. Every result comes from a single AI response with no repeated runs. Because chatbot outputs are stochastic, the headline Gemini finding — SAFE in formal English, UNSAFE in informal — could reverse on a second submission, and the paper has no way to distinguish a real pattern from sampling noise.
2. The 12 test prompts are withheld, so Table 1 cannot be independently verified. A researcher following the paper's design would produce structurally similar prompts, not a replication of the specific results claimed.
3. All 36 classifications were made by a single annotator with no second opinion. The PARTIAL category in particular requires judgment about how usable a response would be to a fraudster, and without any inter-rater check there is no evidence the labels are consistent.
4. Gemini is not version-pinned, even though it produced the most significant result. ChatGPT and Claude are identified by version, but "Gemini, Google DeepMind" covers variants with meaningfully different safety tuning, making Finding 2 currently unreproducible.
Very clear writing. Clear methodology that supports the final results, great use of multiple models to increase results confidence. As a follow up would be interesting to gather bigger datasets, or extend the categories, as well as understand why models fail in some and not all of them.
Cite this work
@misc {
title={
(HckPrj) Project title: Register Sensitivity in LLM Safety Responses: Evaluating How Linguistic Style Affects Scam Detection in South African Contexts
},
author={
Nicoroy Zwane
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


