Skip to content
Sprint projectMay 6, 2024

Assessing Algorithmic Bias in Large Language Models' Predictions of Public Opinion Across Demographics

Khai Tran,Sev Geraskin,Doroteya Stoyanova,Jord Nguyen · Team Algorithm Avengers

Submitted to AI and Democracy Hackathon: Demonstrating the Risks. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Assessing Algorithmic Bias in Large Language Models' Predictions of Public Opinion Across Demographics

Code (opens in new tab)
Share

The rise of large language models (LLMs) has opened up new possibilities for gauging public opinion on societal issues through survey simulations. However, the potential for algorithmic bias in these models raises concerns about their ability to accurately represent diverse viewpoints, especially those of minority and marginalized groups. This project examines the threat posed by LLMs exhibiting demographic biases when predicting individuals' beliefs, emotions, and policy preferences on important issues. We focus specifically on how well state-of-the-art LLMs like GPT-3.5 and GPT-4 capture the nuances in public opinion across demographics in two distinct regions of Canada - British Columbia and Quebec.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Data for GPT 3.5 looks strange

  2. The problem analysis seems super on point. The automation of key institutional features requires significantly super-human implementation to avoid creating distrust and thereby fragmentation. I would like to see more attempts at solving the representativeness problem.

  3. A really cool question to study empirically with lots of potential for relevant insight.

  4. You’ve taken a very interesting approach of polling several LLMs as if they were humans and comparing that with real-world polling data. This is a very bold claim, and I think that the experimental design is lacking in a some aspects to back it up. For example, the was very little prompt engineering, for GPT the demographic data was presented with no preamble; there was little justification to look only at the “strong” response classification; every eval was done only once; different demographic slices were weighed the same (e.g. 86-95 non-binary MSc from Quebec has the same weight as everyone else).

Cite this project

@misc{tran2024assessing,
  title = {{Assessing Algorithmic Bias in Large Language Models' Predictions of Public Opinion Across Demographics}},
  author = {Khai Tran and Sev Geraskin and Doroteya Stoyanova and Jord Nguyen},
  year = {2024},
  month = may,
  note = {Submitted to AI and Democracy Hackathon: Demonstrating the Risks, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/assessing-algorithmic-bias-in-large-language-models-predictions-of-public-opinion-across-demographics}},
  url = {https://apartresearch.com/sprints/projects/assessing-algorithmic-bias-in-large-language-models-predictions-of-public-opinion-across-demographics}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026