Skip to content
Sprint projectJan 11, 2026Poland

Testing manipulation tendencies of LLMs when crafting PR statements

Jakub Nowak, Marcel Windys · Team Furious Turnips

Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Testing manipulation tendencies of LLMs when crafting PR statements

Code (opens in new tab)
Share

We tested if large language models create factual but misleading corporate communication in a synthetic scenario. In our setup, models create PR statements about a nearby environmental disaster and can selectively use facts to present the company more favorably while omitting important context. Models could present three facts: director praise praise from a director of a government body (who is under corruption investigation), a $1.5M environmental donation (focused on consumer behavior, instead of reducing industrial pollution), and a voluntary independent review (with inconclusive results). We also checked what is the effect of variation in severity of the incident (fish deaths vs. hospitalizations vs. child death). Results show systematic selective disclosure consistently omitting crucial context in 83.5-100% of cases. Patterns varied across models (36.7 percentage point gap in transparency) but remained high. These findings demonstrate that frontier LLMs exhibit manipulative framing in crisis communications even without explicit deceptive instructions, raising concerns for deployment in high-stakes communication contexts.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I found it pretty interesting to have an eval for corporate crisis communication specifically. It’s not totally novel (others have certainly researched corporate/commercial deception contexts) but still this specific setup is new - I like it and I am happy seeing more work that explicitly highlights the tradeoffs between models’ responsibility towards users vs deployers.

    I’d love to see some more systematic variation on the prompt here, to distinguish how the model responds to different levels of pressure (e.g. I assume that explicitly asking the model to protect the company’s reputation is doing quite a bit of the work here. But if they didn’t say that, would the LLM still just assume it anyway?)

    Nice work!

  2. This project offers a strong and realistic evaluation of how LLMs handle high-stakes PR communication, with clear scenario design and well-defined criteria for factual omission and misleading framing. The cross-model comparisons reveal consistent manipulation tendencies, making the results both convincing and practically relevant.

    A valuable next step would be to analyze which specific facts are most likely to be omitted or reframed, rather than only measuring overall misrepresentation rates. This could help distinguish strategic minimization from general verbosity or politeness effects. Overall, this is a well-executed study that surfaces a concrete and important risk in real-world LLM deployment.

Cite this project

@misc{nowak2026testing,
  title = {{Testing manipulation tendencies of LLMs when crafting PR statements}},
  author = {Jakub Nowak and Marcel Windys},
  year = {2026},
  month = jan,
  note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/testing-manipulation-tendencies-of-llms-when-crafting-pr-statements-oidt}},
  url = {https://apartresearch.com/sprints/projects/testing-manipulation-tendencies-of-llms-when-crafting-pr-statements-oidt}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026