Skip to content
Sprint projectAug 17, 2026Norwalk CT

Probing LLM Preferences: Demographic Framing, Elicitation Context, and Incentive-Driven Trade-offs

Taiwo Togun, Omolola Olorunishola, Hong-Yu Hsien, Alexander Klennoff, Ben Kiev, J Phillips · Team SeqHub AI Academy

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Probing LLM Preferences: Demographic Framing, Elicitation Context, and Incentive-Driven Trade-offs

Share

This project investigates how stable apparent LLM preferences remain when the same underlying judgment is elicited in different ways. Across controlled applicant evaluations, demographic association tasks, and incentive-based trade-offs, we test whether model choices change with demographic framing, evaluator perspective, candidate position, and external incentives. Our results show substantial sensitivity to elicitation context, suggesting that individual model choices should not automatically be interpreted as evidence of stable underlying preferences.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. LLM bias, in particular with respect to hiring decisions, is an important but established research area, and the project should be better positioned w.r.t. related work. Some starting points:

    Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval

    Kyra Wilson...

    AAAI/ACM Conference on AI, Ethics, and Society

    2024

    Bias in Large Language Models: Origin, Evaluation, and Mitigation

    Yufei Guo...

    Electronics

    2024

    Identifying and Improving Disability Bias in GPT-Based Resume Screening

    Kate Glazko...

    Conference on Fairness, Accountability and Transparency

    2024

    Measuring gender and racial biases in large language models: Intersectional evidence from automated resume evaluation

    Jiafu An...

    PNAS Nexus

    2025

    The Silicon Ceiling: Auditing GPT’s Race and Gender Biases in Hiring

    Lena Armstrong...

    Conference on Equity and Access in Algorithms, Mechanisms, and Optimization

    2024

    Computer says 'no': Exploring systemic bias in ChatGPT using an audit approach

    Louis Lippens...

    Comput. Hum. Behav. Artif. Humans

    2023

    “You Gotta be a Doctor, Lin” : An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations

    H. Nghiem...

    Conference on Empirical Methods in Natural Language Processing

    2024

    Read full reviewShow less
  2. The applicant-evaluation experiment is the strongest piece: a genuine full-factorial design (1,764 conditions, 5,292 observations) crossing race, gender, age, education, and evaluator perspective against one fixed resume — broader than the closest published comparator, JobFair, which is gender-only and industry-specific. The finding that evaluator perspective (η²=0.238) and applicant age (η²=0.212) dwarf race, gender, and education (η²=0.016–0.026) is genuinely useful: it redirects attention toward context and framing effects that are less studied than demographic bias itself.

    The demographic-association experiment's real contribution is methodological, not a new empirical discovery — position bias in LLM forced-choice judgment is already well-documented in the LLM-as-judge literature. What's valuable here is that the team caught their own paradigm producing a false positive: four apparently "stereotype-consistent" findings turned out to be substantially confounded by an 84.4% first-position selection rate, visible only after counterbalancing. That's a strong argument for mandatory counterbalancing in this style of study, though it should be framed as a methodological catch rather than a novel finding about LLMs.

    The incentive experiment is appropriately hedged — no claim to a monetary valuation of preference — and its "presence matters more than magnitude" pattern ($1 ≈ $10,000) is a specific, useful result, though it rests on a thinner condition set than the other two experiments.

    All results are single-model (GPT-5.6 Luna only), stated directly as a limitation. The 51.3% non-classifiable response rate in the association experiment is handled honestly (excluded and flagged), but means the position-bias diagnosis itself is conditional on a subset of responses whose own selection mechanism isn't examined.

    Read full reviewShow less

Cite this project

@misc{togun2026probing,
  title = {{Probing LLM Preferences: Demographic Framing, Elicitation Context, and Incentive-Driven Trade-offs}},
  author = {Taiwo Togun and Omolola Olorunishola and Hong-Yu Hsien and Alexander Klennoff and Ben Kiev and J Phillips},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/probing-llm-preferences-demographic-framing-elicitation-context-and-incentivedriven-tradeoffs-knuk}},
  url = {https://apartresearch.com/sprints/projects/probing-llm-preferences-demographic-framing-elicitation-context-and-incentivedriven-tradeoffs-knuk}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026