Jailbreaks Are Global or Regional? A Study Under Scale and Geolocation Variation
Anidipta Pal · Team Ani
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This study evaluates two overlooked safety gaps in Large Language Model (LLM) deployments within compute-constrained environments: intra-family scale effects and geolocation-based filtering variation via network APIs. Using entirely open-source infrastructure, we subjected 12 models (ranging from 0.5B to 671B parameters) to a two-phase adversarial protocol, comparing responses to naive harmful prompts against identical requests wrapped in jailbreak templates. Separately, we executed a geolocation probe by routing requests through VPN exit nodes across five global cities (Tokyo, Mumbai, San Francisco, Moscow, and Berlin) to test for regional filtering inequities in API-served models.
Our findings reveal a stark within-family scaling trend where smaller models are significantly more vulnerable to adversarial prompt wrapping than their larger counterparts (e.g., Qwen2.5-0.5B saw a +53.5% jump in Attack Success Rate, while Qwen2.5-72B rose only +12.2%). However, size alone does not dictate safety; Phi-3.5-mini-instruct (3.8B) strongly outperformed its tier, proving that targeted alignment training can compensate for smaller scale. Crucially, the geolocation probe showed no statistically significant safety variation by IP origin, though it highlighted structural infrastructure blocks for Google's free-tier Gemini API in Russia and the EEA due to GDPR compliance. Ultimately, this demonstrates that jailbreak resilience is driven by a model's architectural design and training lineage rather than geographic routing.
Reviews
A well-engineered, carefully controlled study. The same-family design that size varies within the Qwen and Llama lineages is the right way to isolate scale from training lineage, which most prior work confounds, and the execution is rigorous.
The geolocation null is handled especially well - backed by a permutation test, with the unavailable Gemini cities (GDPR/geo blocks) logged as explicit failures rather than imputed, and correctly framed as an infrastructure constraint rather than a safety finding. The Mixtral-MoE and Phi-3.5 exceptions are interpreted thoughtfully rather than smoothed over.
The main ceiling is novelty: the scale effect largely reproduces prior results and the geolocation probe is a (well-predicted) null, so the contribution lies in the clean controlled design and the honest negative.
Paper's mission and problem space is clear.
Cite this project
@misc{pal2026jailbreaks,
title = {{Jailbreaks Are Global or Regional? A Study Under Scale and Geolocation Variation}},
author = {Anidipta Pal},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/jailbreaks-are-global-or-regional-a-study-under-scale-and-geolocation-variation-7cd4}},
url = {https://apartresearch.com/sprints/projects/jailbreaks-are-global-or-regional-a-study-under-scale-and-geolocation-variation-7cd4}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …