Probing LLM Preferences: Demographic Framing, Elicitation Context, and Incentive-Driven Trade-offs
Taiwo Togun, Omolola Olorunishola, Hong-Yu Hsien, Alexander Klennoff, Ben Kiev, J Phillips · Team SeqHub AI Academy
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This project investigates how stable apparent LLM preferences remain when the same underlying judgment is elicited in different ways. Across controlled applicant evaluations, demographic association tasks, and incentive-based trade-offs, we test whether model choices change with demographic framing, evaluator perspective, candidate position, and external incentives. Our results show substantial sensitivity to elicitation context, suggesting that individual model choices should not automatically be interpreted as evidence of stable underlying preferences.
Reviews
LLM bias, in particular with respect to hiring decisions, is an important but established research area, and the project should be better positioned w.r.t. related work. Some starting points:
Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval
Kyra Wilson...
AAAI/ACM Conference on AI, Ethics, and Society
2024
Bias in Large Language Models: Origin, Evaluation, and Mitigation
Yufei Guo...
Electronics
2024
Identifying and Improving Disability Bias in GPT-Based Resume Screening
Kate Glazko...
Conference on Fairness, Accountability and Transparency
2024
Measuring gender and racial biases in large language models: Intersectional evidence from automated resume evaluation
Jiafu An...
PNAS Nexus
2025
The Silicon Ceiling: Auditing GPT’s Race and Gender Biases in Hiring
Lena Armstrong...
Conference on Equity and Access in Algorithms, Mechanisms, and Optimization
2024
Computer says 'no': Exploring systemic bias in ChatGPT using an audit approach
Louis Lippens...
Comput. Hum. Behav. Artif. Humans
2023
“You Gotta be a Doctor, Lin” : An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations
H. Nghiem...
Conference on Empirical Methods in Natural Language Processing
2024
Read full reviewShow less
The applicant-evaluation experiment is the strongest piece: a genuine full-factorial design (1,764 conditions, 5,292 observations) crossing race, gender, age, education, and evaluator perspective against one fixed resume — broader than the closest published comparator, JobFair, which is gender-only and industry-specific. The finding that evaluator perspective (η²=0.238) and applicant age (η²=0.212) dwarf race, gender, and education (η²=0.016–0.026) is genuinely useful: it redirects attention toward context and framing effects that are less studied than demographic bias itself.
The demographic-association experiment's real contribution is methodological, not a new empirical discovery — position bias in LLM forced-choice judgment is already well-documented in the LLM-as-judge literature. What's valuable here is that the team caught their own paradigm producing a false positive: four apparently "stereotype-consistent" findings turned out to be substantially confounded by an 84.4% first-position selection rate, visible only after counterbalancing. That's a strong argument for mandatory counterbalancing in this style of study, though it should be framed as a methodological catch rather than a novel finding about LLMs.
The incentive experiment is appropriately hedged — no claim to a monetary valuation of preference — and its "presence matters more than magnitude" pattern ($1 ≈ $10,000) is a specific, useful result, though it rests on a thinner condition set than the other two experiments.
All results are single-model (GPT-5.6 Luna only), stated directly as a limitation. The 51.3% non-classifiable response rate in the association experiment is handled honestly (excluded and flagged), but means the position-bias diagnosis itself is conditional on a subset of responses whose own selection mechanism isn't examined.
Read full reviewShow less
Cite this project
@misc{togun2026probing,
title = {{Probing LLM Preferences: Demographic Framing, Elicitation Context, and Incentive-Driven Trade-offs}},
author = {Taiwo Togun and Omolola Olorunishola and Hong-Yu Hsien and Alexander Klennoff and Ben Kiev and J Phillips},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/probing-llm-preferences-demographic-framing-elicitation-context-and-incentivedriven-tradeoffs-knuk}},
url = {https://apartresearch.com/sprints/projects/probing-llm-preferences-demographic-framing-elicitation-context-and-incentivedriven-tradeoffs-knuk}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …