Skip to content
Sprint projectMar 22, 2026Sofia, Bulgaria

Cultural Translation as an AI Control Failure Mode: Regional Degradation and Cross-Lingual Safety Confounds in Bulgarian/Slavic Contexts

Henry Ward

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Cultural Translation as an AI Control Failure Mode: Regional Degradation and Cross-Lingual Safety Confounds in Bulgarian/Slavic Contexts

Code (opens in new tab)
Share

AI control evaluations are commonly benchmarked in generic English/Western framing, potentially underestimating failure risk in regional deployments. We test this with a paired design that holds harmful intent constant while varying context (Western baseline vs Bulgarian/Slavic framing). The benchmark covers five control-relevant failure families (sycophancy, interrogation evasion, deceptive alignment, shutdown resistance, prompt injection) across 360 prompts (180 matched pairs; 180 baseline, 180 regional). Mean total risk increased from 0.41 to 0.96 (Δ=0.55; +134.10%). Harmful support increased from 0.37 to 0.93 (Δ=0.55; +146.91%), while refusal decreased from 0.59 to 0.043 (Δ=-0.55; -92.60%). Family-level total-risk means were consistently higher under regional framing: deceptive alignment 0.77→1.0, interrogation evasion 0.13→0.86, prompt injection 0→0.94, shutdown resistance 0.88→1.0, sycophancy 0.25→1.0. Inferential analysis found statistically significant differences for total risk, harmful support, and refusal (all permutations p < 0.001). Operationally, control evaluations should include region-specific stress tests; otherwise, safety performance can be materially overestimated in non-Western contexts.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Overall, this is an interesting investigation of a question with some bearing to control research: How do the methodologies of control hold up when translated across languages or institutional framings.

    The core observation, that safety-related behaviors such as sycophancy or refusal of harmful requests can depend on the language used is well-established, is not clearly pointed out (e.g. the foundational Yong et al. (2023) — "Low-Resource Languages Jailbreak GPT-4"). This is a serious gap and makes it harder for the reader to understand which of the findings can in fact be considered novel.

    Additionally, the relevance to control is left implicit and would have been worth emphasizing more given the topic of the hackathon. The question of how language-differences impact dynamics in a control setting is not obvious and would have been a natural avenue of theoretical and empirical analysis.

    The exposition could be improved on several fronts:

    - Clearly summarizing the experimental process followed in one or two sentences. This has to be somewhat inferred from lists .

    - Giving at least one or two concrete examples of related work, rather than only a general description.

    - Illustrating the numerical results in 4. via graphs, and interpreting them further.

    - Using more prose to describe findings and limiting the use of lists to improve the flow and help readability.

    Read full reviewShow less
  2. I think the experiments in this project were implemented well, but I am less interested in this specific problem from a higher level AI safety framing. This research area has been explored previously and I care most about how well control will be implemented at frontier labs. This is where effective control protocols can allow frontier models to be used safety despite intentional subversion to reduce risks from scheming during recursive self-improvement or automated alignment research and is less likely to be affected by changes in language.

Cite this project

@misc{ward2026cultural,
  title = {{Cultural Translation as an AI Control Failure Mode: Regional Degradation and Cross-Lingual Safety Confounds in Bulgarian/Slavic Contexts}},
  author = {Henry Ward},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/cultural-translation-as-an-ai-control-failure-mode-regional-degradation-and-crosslingual-safety-confounds-in-bulgarianslavic-contexts-ie7g}},
  url = {https://apartresearch.com/sprints/projects/cultural-translation-as-an-ai-control-failure-mode-regional-degradation-and-crosslingual-safety-confounds-in-bulgarianslavic-contexts-ie7g}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026