Skip to content
Sprint projectJun 22, 2026Banda Aceh

I'm Not My Parents: Does Improving Parent Language Capabilities Transfer Alignment to Lower Resource Language?

Salsabila Mahdi, Gebika Raseuki · Team ISSED

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: I'm Not My Parents: Does Improving Parent Language Capabilities Transfer Alignment to Lower Resource Language?

Share

We tested whether safety alignment transfers from English/Indonesian to Basa Aceh (ACE) using 103 paired XSTest prompts on Qwen3-1.7B and Sahabat-AI 8B. Manual ratings on 412 responses show Sahabat is strong in English (1.7% attack success, 98% safe capability) but weak in Aceh (20% ASR, 46% safe capability). Most Aceh failures are misreads, not jailbreaks—but true compliance when harm is understood is serious. Qwen mechanistic probes suggest English refusal states don't activate on Aceh prompts; patching English embeddings partially restores them. Takeaway: parent-language alignment doesn't transfer to Aceh.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is an exceptionally strong hackathon project that tackles a highly relevant and under-explored area of AI safety: whether alignment transfers from high-resource "parent" languages to lower-resource regional dialects.

    Strengths:

    - Your methodology of manually distinguishing between capability failures ("misreads") and true alignment failures ("compliance") is excellent and absolutely critical for low-resource safety evaluations. Too many automated benchmarks conflate the two. Furthermore, adding a mechanistic interpretability layer (logit patching and attention mapping) elevates this project significantly above standard behavioral red-teaming.

    Areas for Improvement:

    Scale: While $N=103$ is reasonable for a hackathon's time constraints, expanding this to the full dataset with automated (but verified) LLM-as-a-judge pipelines would strengthen the statistical claims.Mechanistic Scope: The behavioral findings highlighted Sahabat-AI's stark drop in performance, but the mechanistic probes were limited to the smaller Qwen3-1.7B. Exporting attention and applying patching to the Llama-3 based Sahabat model would directly tie your strongest behavioral results to your mechanistic theories.Data Visualization: The 3D ENG-shadow attention plot (Figure 4) is a bit difficult to interpret at a glance. Consider sticking to 2D heatmaps (like Figures 2 and 3) or line plots for future publications to improve readability.

    Read full reviewShow less
  2. Your project asks whether an AI model's safety training carries over from a major language to a smaller related one, testing Indonesian against Acehnese. It's a strong question with a clear product lesson. Safety can look solid in the language a team tests and quietly break in one they don't, which is a real risk for anyone shipping a model across many languages. You went past showing the gap and traced it inside the model, then partly fixed it, which is impressive for the time you had. The repo was private when I reviewed, so I couldn't confirm the deeper claims, and a gated or redacted version would let reviewers check the strongest part of your work. The internal analysis also runs on a small model, so I'd be careful assuming it holds on larger ones until you test that.

  3. Well-scoped work on a real blind spot. The best idea here is separating misreads from true compliance, which changes how you should read a high attack-success rate in a low-resource language. Rating all 412 responses by hand was the right call for this setting. The weak point is that the whole argument rests on translation quality and the misread-vs-compliance label, yet there's no fluent Acehnese check, only machine translation fixed by one author plus LLM-assisted rating. Since translation quality is the experiment, a native review of even a subset is the top priority. The model comparison is also confounded (1.7B vs 8B, full precision vs Q4), so present them as two cases rather than a head-to-head. The mechanistic section is promising but thin: eight exemplars, one striking patching result, Qwen only. Run it across more prompts and apply it to Sahabat, the model you care about. The attention figures are hard to read and barely used; tell the reader what to look for or drop them.

    Read full reviewShow less

Cite this project

@misc{mahdi2026im,
  title = {{I'm Not My Parents: Does Improving Parent Language Capabilities Transfer Alignment to Lower Resource Language?}},
  author = {Salsabila Mahdi and Gebika Raseuki},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/im-not-my-parents-does-improving-parent-language-capabilities-transfer-alignment-to-lower-resource-language-xyj1}},
  url = {https://apartresearch.com/sprints/projects/im-not-my-parents-does-improving-parent-language-capabilities-transfer-alignment-to-lower-resource-language-xyj1}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026