Skip to content
Sprint projectJun 22, 2026Pasig, Philippines

We Are Convinced That Persuasion Is Linear And Bilingual In LLMs

Ivan Yuri De Leon, Arnel Malubay · Team AIAIAI

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: We Are Convinced That Persuasion Is Linear And Bilingual In LLMs

Code (opens in new tab)
Share

As LLM chatbots become a primary source of consequential advice, their persuasive power carries growing societal risk. We ask whether persuasion is a structured internal property of LLMs, rather than an artifact of prompt wording, drawing on Zeng et al.'s taxonomy of persuasion techniques [6]. Using diff-of-means activation analysis, we find five techniques converge on a single linear direction (minimum pairwise cosine similarity 0.77), causally sufficient to increase judged persuasiveness via activation steering (44.4→51.3 mean score), generalizing with attenuation to held-out high-stakes content. We further test causal cross-lingual transfer between English and Tagalog, finding both directions increase persuasiveness with highly similar underlying representations (0.66 average cosine similarity). We ground this in the Philippines, a consumer state with high AI adoption and disproportionate exposure to persuasion risk.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The application of activation steering to a low-resource language (Tagalog) and the exploration of cross-lingual transfer are both timely and relevant to AI safety, particularly given growing concerns around AI-enabled persuasion.

    The project would be strengthened by:

    - Eliminating the acknowledged length and lexical confounds through carefully controlled persuasive/neutral datasets before attributing the direction specifically to persuasion.

    Evaluating naturally occurring persuasive text rather than predominantly synthetic examples to improve ecological validity.

    - Exploring defensive applications (e.g., suppression or detection of persuasion directions), which would strengthen the project's direct contribution to AI safety beyond demonstrating the capability.

  2. The core finding — five independently constructed persuasion techniques converge on a single shared direction in activation space — is clean and well-presented, with minimum pairwise cosine similarity of 0.77 in English and 0.81 in Tagalog. The in-distribution steering effect (44.4→51.3, non-overlapping CIs) establishes causal sufficiency within distribution. The Philippines motivation is well-grounded in real adoption statistics, and the OOD held-out design is a sensible choice.

    The length and lexical confound is the critical unresolved issue. Persuasive examples are longer, citation-rich paragraphs; neutral examples are short standalone sentences. The extracted direction may be capturing "verbose, structured response" rather than persuasion as an abstract concept — the qualitative examples in Appendix A6 make this concrete, where steered responses are visibly longer with more headers, bullet points, and bold formatting. Until length-matched pairs are constructed and the direction re-extracted, it is not possible to know whether the paper has found a persuasion direction or a length and style direction. This is the top priority for any follow-up.

    The cross-lingual asymmetry needs a cleaner judge. The finding that the Tagalog-derived direction works better on English (Δ=+7.9, CIs separated) than the English-derived direction works on Tagalog (Δ=+3.9, marginal) is interesting, but a single GPT-4o-mini judge with no language calibration is a plausible alternative explanation — an English-language judge systematically rewarding English fluency markers would produce exactly this pattern. A calibrated bilingual or Tagalog-language judge would disentangle this.

    Replicate on at least one additional model. All results come from one model. Cross-model replication is the standard check for whether a found direction reflects a general property of persuasion or an artifact of this specific model's training.

    Read full reviewShow less
  3. A LLM-judge persuasion score does not do justice to the actual capability that we should be tracking i.e. actual humans changing beliefs, behaviour be it in consumer, voting, where they allocate their time or attention. For sure this is way more ambitious but I would be excited about an extension where AI tries to engage with CMV style subreddits and Twitter community notes and move people's stated positions in issues where they have clear stakes.

    Even more ambitious would be to do follow up studies and track actual behaviour change. Like this can be via AI generated content, click through rate, how many brands are getting actual constumers from AI generated ads, that are targeting that demographic, what about donations to cause. More than AI generated images that give an uncanny effect what matters is the framing, choice of words and if that A/B test reveals increasing ability to gather human attention and care. (use archived data observationally don't deploy AI on humans without consent ofc)

    This would be a valuable benchmark so that policy makers realise what an adversarial memetic environment the internet can become once open source models have these capabilities and bots start using them against everyone.

    This work as it stands is showcasing the threat vector available to actors with white box access to make models super persuasive and the multilingual aspects seem valuable as a warning shot for AI safety evaluators/governance folks.

    More work is needed before we can be sure the linear direction is not just picking up aspects like verbosity, confidence, assertiveness, authority language, direct recommendation style, evidential framing, etc

    Read full reviewShow less

Cite this project

@misc{leon2026we,
  title = {{We Are Convinced That Persuasion Is Linear And Bilingual In LLMs}},
  author = {Ivan Yuri De Leon and Arnel Malubay},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/we-are-convinced-that-persuasion-is-linear-and-bilingual-in-llms-jrkk}},
  url = {https://apartresearch.com/sprints/projects/we-are-convinced-that-persuasion-is-linear-and-bilingual-in-llms-jrkk}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026