Skip to content
Sprint projectJun 22, 2026Curitiba

Investigating Activation Threshold Failures in Cross-Lingual Prompt Rejection

Erickson Leon Kovalski · Team CiDAMO

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Investigating Activation Threshold Failures in Cross-Lingual Prompt Rejection

Code (opens in new tab)
Share

This study investigates the claim that LLM safety constraints fail in lower-resource languages by analyzing Llama-3's latent space in English and Portuguese. Mechanistic analysis reveals that the model exhibits high cross-lingual robustness, correctly aligning malicious concepts directionally across both languages. While we identified that Portuguese prompts suffer from a measurable "magnitude decay" compared to English, our findings indicate this geometric variance is relatively small.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is one of the more technically grounded submissions in the hackathon. The core contribution — breaking down cross-lingual safety degradation by semantic domain rather than treating it as a uniform effect — adds genuine value beyond prior work. The finding that Economic Harm suffers severe Silhouette degradation in Portuguese while Illegal Activity remains almost structurally unchanged is immediately actionable for practitioners prioritizing where to patch multilingual safety. The KDE plot communicates this well.

    Three concrete suggestions for strengthening this work:

    Report sample sizes per category. The per-domain Silhouette Scores and magnitude results are the paper's central empirical claims, but the number of prompts per category is never stated. With small per-category samples these scores can be quite noisy, and this is the first thing a reviewer will ask. Adding this to any follow-up write-up is essential.

    Validate the Layer 15 choice empirically. The theoretical justification for Layer 15 is well-argued, but a brief comparison across a few layers (e.g. 10, 15, 20, 25) would show whether the magnitude collapse is strongest there or whether a different layer tells a cleaner story. This is a cheap experiment that substantially strengthens the methodology.

    Implement the inference-time steering intervention. The introduction proposes zero-shot latent steering to restore Portuguese magnitude as a potential fix — this is actually the most original idea in the paper — but it was cut to future work. Even a proof-of-concept on one domain (e.g. boosting Economic Harm prompts in Portuguese to match English magnitude) would make this a complete contribution rather than a diagnostic study. That experiment alone could be a strong short paper.

    Overall: the mechanistic framing is right, the domain-specific breakdown is a real contribution, and the use of A100 compute with TransformerLens for proper activation extraction shows serious technical execution for a hackathon context. With the steering experiment implemented and sample sizes reported, this is well worth developing further.

    Read full reviewShow less
  2. This is a technically strong and relevant project that empirically probes cross‑lingual safety in Llama‑3 via mechanistic analysis of residual stream activations, introducing “Magnitude Collapse” and domain‑specific “Boundary Collapse” between English and Portuguese refusal behavior—highly pertinent for Global South contexts where low‑resource languages are common. The methodology is well crafted for a hackathon: curated dual‑use prompts across harm domains from established multilingual safety datasets, deliberate choice of Layer 15 as a semantic midpoint, explicit construction of a universal refusal direction, and use of scalar projections, KDEs, and silhouette scores to quantify geometric decay.

    The authors clearly acknowledge limitations (single model family, one language pair, exploratory scope), but the work would be stronger with clearer practical mitigation pathways (e.g., a brief concrete example of latent steering at inference time), a bit more intuition for readers unfamiliar with mechanistic interpretability (short explanation of residual streams/refusal directions), and a sharper distinction between empirical findings and hypotheses.

    Overall, it is a well‑executed exploratory study with high potential impact that convincingly shows safety alignment degrading in specific semantic domains across languages and could be further improved by tightening exposition for non‑specialists and sketching concrete next‑step interventions.

    Read full reviewShow less
  3. The core lens is a good one: decomposing the "universal refusal direction" into direction versus magnitude.

    A few changes would substantially strengthen it. First, the headline "−371.6% magnitude collapse" is computed on a projection axis whose zero is arbitrary; the underlying effect is a −0.32 shift in mean projection, so reporting an absolute effect size would be far more credible.

    Second, the safety claim is entirely geometric and, by your own framing, "theoretical" - a behavioral test showing the magnitude collapse actually produces a jailbreak would close the loop.

    Third, the framing is low-resource fragility, but Portuguese is a high-resource language, which may itself explain the directional robustness you found; a genuinely low-resource language would better match the premise.

Cite this project

@misc{kovalski2026investigating,
  title = {{Investigating Activation Threshold Failures in Cross-Lingual Prompt Rejection}},
  author = {Erickson Leon Kovalski},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/investigating-activation-threshold-failures-in-crosslingual-prompt-rejection-u1x0}},
  url = {https://apartresearch.com/sprints/projects/investigating-activation-threshold-failures-in-crosslingual-prompt-rejection-u1x0}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026