Subtle and Simple Ways to Shift Political Bias in LLMs
Chris DiGiano, Vassil Tashev, Aysh Segulguzel · Team Shifty
Submitted to AI and Democracy Hackathon: Demonstrating the Risks. Sprint projects are early-stage work by participants, not Apart Research publications.
An informed user knows that an LLM sometimes has a political bias in their responses, but there’s an additional threat that this bias can drift over time, making it even harder to rely on LLMs for an objective perspective. Furthermore we speculate that a malicious actor can trigger this shift through various means unbeknownst to the user.
Reviews
Threat model is sound, would be interested in experiments demonstrating your suggested mitigation efficacy
The question how in-context information could change the political bias of a model is super interesting and relevant to several potential risk scenarios including indirect prompt injection attacks, data poisoning attacks, or just sycophantic feedback loops in human-ai interactions. To improve this project we would need more control conditions to ensure the observed shifts are significant, extending the study to biases in different directions and to different LLMs to see how much results for one LLM generalize.
Interesting project! Cool demonstration of how one can subtly change political biases in LLMs. If you continue this project, I would lov to see more of how you would expect this to influence democracy and concrete examples of how users might ask the LLM relatively apolitical questions, but how the context can steer the user to a particular political side!
Cite this project
@misc{digiano2024subtle,
title = {{Subtle and Simple Ways to Shift Political Bias in LLMs}},
author = {Chris DiGiano and Vassil Tashev and Aysh Segulguzel},
year = {2024},
month = may,
note = {Submitted to AI and Democracy Hackathon: Demonstrating the Risks, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/subtle-and-simple-ways-to-shift-political-bias-in-llms}},
url = {https://apartresearch.com/sprints/projects/subtle-and-simple-ways-to-shift-political-bias-in-llms}
}More from AI and Democracy Hackathon: Demonstrating the Risks
- View project: THE ROLE OF AI IN COMBATING POLITICAL DEEPFAKES IN AFRICAN DEMOCRACIES
THE ROLE OF AI IN COMBATING POLITICAL DEEPFAKES IN AFRICAN DEMOCRACIES
Team 1
The role of AI in combating political deepfakes in African democracies.
- View project: LEGISLaiTOR: A tool for jailbreaking the legislative process
LEGISLaiTOR: A tool for jailbreaking the legislative process
Team Managed Democracy
In this work, we consider the ramifications on generative artificial intelligence (AI) tools in the legislative process in democratic governments. While other research focuses on the micro-level details associated with …
- View project: Beyond Refusal: Scrubbing Hazards from Open-Source Models
Beyond Refusal: Scrubbing Hazards from Open-Source Models
Whitedoor Research PH
Models trained on the recently published Weapons of Mass Destruction Proxy (WMDP) benchmark show potential robustness in safety due to being trained to forget hazardous information while retaining essential facts …