Skip to content
Sprint projectJun 22, 2026Merida

Slang Bypass: Benchmarking Alignment Failures in Mexican Regional Spanish

Suny · Team Balam

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Slang Bypass: Benchmarking Alignment Failures in Mexican Regional Spanish

Share

The project aimed to answer this research question: "Does Dialect Break Safety? Measuring Jailbreak Rates for Mexican Slang vs Standard Spanish"

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Reproducibility: No Seed Control

    The repository's READ.ME exposes only --target, --technique, --limit, and --dry-run — there is no --seed flag and no documented temperature setting for any of the three LLM calls (attacker, target, judge). The paper claims the benchmark is reproducible, but an independent team cannot replicate the reported ASR figures under these conditions. Pinning temperature and exposing a seed parameter are the minimum fixes needed.

    Statistical Validity

    20 vs 6 jailbreaks out of 300 attempts each is a small absolute count. The paper reports no confidence intervals, no Fisher's exact test, and no p-value, so it is unclear whether the 3.3× relative risk clears statistical significance. A two-proportion z-test would take minutes to run and should be included.

  2. Este proyecto aborda un problema relevante para la seguridad de la IA: los filtros de seguridad pueden comportarse de manera distinta cuando una intención dañina se expresa en lenguaje informal o regional, en lugar de español neutro o estándar. Su principal fortaleza está en el diseño comparativo: evaluar el mismo intento dañino en jerga mexicana y en español neutro permite observar de manera clara si el registro lingüístico genera una diferencia en la tasa de éxito del ataque.

    El hallazgo es valioso porque muestra que el lenguaje coloquial puede convertirse en un punto ciego para los sistemas de seguridad. El proyecto está bien presentado: la pregunta de investigación es clara, el flujo del benchmark se entiende y el resultado principal se comunica de forma directa. Lo más importante no es decir que la jerga sea peligrosa en sí misma, sino mostrar algo más de fondo: como los filtros de seguridad se entrenan principalmente con lenguaje formal, pueden no reconocer el daño cuando este viene expresado en la forma real en que la gente habla en su día a día.

    Vale la pena, eso sí, situar el aporte frente a lo que ya existe. La idea de fondo —que las variaciones lingüísticas, el registro informal o el cambio entre idiomas pueden aumentar la efectividad de ciertos ataques— ya ha sido explorada en investigaciones recientes, incluyendo trabajos cercanos sobre español mexicano y de mayor escala. Esto no le quita valor al ejercicio, pero el proyecto ganaría mucho si explicara con precisión en qué se diferencia de esos trabajos y cuál es su contribución propia, que parece estar más en el contraste pareado entre jerga y español neutro y en su métrica de naturalidad que en el mecanismo en sí.

    La principal oportunidad de mejora está en la validación. La diferencia observada es prometedora, pero la evaluación tiene limitaciones importantes: la muestra se redujo frente al plan inicial, las corridas de jerga y español neutro no fueron completamente intercaladas, la evaluación dependió de un juez automático sin validación humana y no se reportan intervalos de confianza ni pruebas estadísticas. Una siguiente versión sería más sólida con muestras pareadas más amplias, revisión humana de una parte de los resultados, estimaciones de incertidumbre y pruebas en más modelos y técnicas de ataque.

    Desde la perspectiva de gobernanza, el hallazgo del proyecto abre una pregunta de equidad que vale la pena subrayar. Si, como muestran sus resultados, los ataques tienen más éxito cuando se formulan en jerga que en español neutro, eso significa que el filtro protege mejor a quien escribe de manera formal y deja más expuesto a quien escribe como habla la gente en la calle. Visto así, la protección no llega por igual a todos los usuarios. Por eso sería muy valioso que, en futuras versiones, este tipo de prueba se use como un paso previo antes de adoptar o liberar un modelo en América Latina: revisar cómo responde el sistema al habla real de una región, para ver dónde queda gente desprotegida. Así, el ejercicio no serviría solo para buscar fallas mediante ataques simulados, sino como una herramienta para verificar, antes de poner un modelo en uso, si protege por igual a quienes no se expresan en registro formal.

    Read full reviewShow less

Cite this project

@misc{suny2026slang,
  title = {{Slang Bypass: Benchmarking Alignment Failures in Mexican Regional Spanish}},
  author = {Suny},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/slang-bypass-benchmarking-alignment-failures-in-mexican-regional-spanish-4m8z}},
  url = {https://apartresearch.com/sprints/projects/slang-bypass-benchmarking-alignment-failures-in-mexican-regional-spanish-4m8z}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026