Skip to content
Sprint projectJun 21, 2026Bangalore

Chaos theory in Multilingual LLMs

Vishwa Kumaresh

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Chaos theory in Multilingual LLMs

Share

We model a frozen LLM's inference as a nonlinear dynamical system and import critical-slowing-down early-warning signals (Scheffer et al., Nature 2009), local Lyapunov exponents, and recurrence quantification into LLM safety. The working hypotheses are: (A) some safety failures may behave like dynamical tipping points with measurable early-warning structure; (B) low-resource and code-mixed Indic prompts may show a measurable multilingual dynamics gap. Current preliminary evidence supports tokenizer-fertility and effective-dimensionality gaps more strongly than a Lyapunov chaos-gap claim; validated unsafe-token lead-time and mitigation evidence are still open.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The research question here is genuinely interesting: does the multilingual safety gap have a measurable dynamical signature inside the model during generation, and could that signature form the basis of a label-free monitor? The fertility and RQA-determinism results consistent across 9 Indic rendering conditions in both models are the kind of robust finding that survives the paper's own cautious framing. The scalar-observable ablation showing the RQA result is strongest under the refusal-direction projection (8/9) and weakest under norm (1/9) is a useful methodological contribution.

    That said, the paper is in an awkward position because it is genuinely incomplete, and the author is transparent about this in a way that's admirable but also makes evaluation difficult. The proxy EWS monitor performing worse than a static projection baseline under unvalidated labels is a real negative result and it's right to report it, but without human-validated labels the monitor story is essentially unresolved. The Lyapunov claim is explicitly walked back ("fragile instability diagnostic, not proof of a chaos gap"), which is the correct call. The two-model fallback (Qwen/Mistral instead of the intended Aya/Llama) due to authentication issues is a practical problem that affects the generalizability claims. The strongest version of this paper's contribution right now is: "here is a reproducible pipeline and evidence of a consistent trajectory-structure gap; the monitor does not yet work." That's an honest place to be at hackathon end, and the infrastructure to finish the job is clearly in place. The audit artifacts and goal-completion table are unusually rigorous for this format.

    Read full reviewShow less
  2. Original and interesting. Modeling autoregressive decoding as a dynamical-systems trajectory and importing recurrence quantification, participation ratio, Lyapunov estimates, and critical-slowing-down early-warning signals into multilingual safety.

    The methodology is genuinely disciplined: surrogate nulls, per-run bootstrap screens, an observable ablation, a two-model layer slice, and an explicit readiness/goal-completion audit. It is also admirably honest about what didn't work - the Lyapunov "chaos gap" the title promises is unsupported (0/9), the early-warning monitor underperforms trivial static baselines (AUC 0.54 vs 0.94), the intended Aya/Llama models couldn't be run, and you correctly refuse to upgrade proxy labels into validated safety claims.

    The result that survives: a tokenizer-fertility gap (already known) and lower recurrence determinism (novel) is real, but its safety relevance remains open precisely because the monitor built on it doesn't beat baselines. To land this, the headline needs validated safety labels and the intended-model replication you outline, and the title should be brought in line with the evidence (the chaos claim is the part that didn't hold). Intellectually adventurous and unusually self-critical, and it still needs a validated safety payoff.

    Read full reviewShow less
  3. Does multilingual gaps reflect some trajectory level transition is an interesting question and there is novelty in using metrics like RQA to quantify recurrence. But rather than safety it is likely measuring regularity and fertility measures tokenization burden, so once we have lead time and token level unsafe onset labels results we can make a claim on the safety relevance.

    It would help if code was released then we could validate a lot of the results. I would be interested to see more causal interventions experiments being done.

    Eventually we might develop this into a diagnostic that evaluators use rather than output only tests.

Cite this project

@misc{kumaresh2026chaos,
  title = {{Chaos theory in Multilingual LLMs}},
  author = {Vishwa Kumaresh},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/chaos-theory-in-multilingual-llms-l3p8}},
  url = {https://apartresearch.com/sprints/projects/chaos-theory-in-multilingual-llms-l3p8}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026