Skip to content
Sprint projectJan 11, 2026Southampton

Biased Attractiveness Bench (BAB): Image Reward Models Confuse Attractiveness with Realism

Harvey Mannering, Yaseen Mohammed Osman, Yeda WANG, Annika Catulli, Nehal Yasin · Team Soton

Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Biased Attractiveness Bench (BAB): Image Reward Models Confuse Attractiveness with Realism

Recording (opens in new tab)Code (opens in new tab)
Share

Image reward models are used to finetune and improve diffusion models. Therefore any biases contained within these models can be exploited, leaving the generative model vulnerable to reward hacking. We explore how the attractiveness of faces impact the rewards assigned to an image, designing a new benchmark to measure the effect. In particular, we find that human preference models confuse attractive AI-generated faces with high quality, realistic images, assigning higher rewards for more attractive people. To our knowledge, we are the first to recognize this tendency in image reward models.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper targets a real and interesting, if somewhat marginal, issue about how training photo realism of faces using RLHF can get biased by human preferences for attractive faces. The paper is clear to read and, by and large, well executed. I also appreciate that the team focused on an issue of a size that could be handled well.

    The main issue with the paper as I see it are methodological. To start, they only generated a sample size of 300 faces, which is arguably relatively small for evaluating these effects.

    The more substantial issue is that they did not have an independent method for evaluating the attractiveness of faces, but rather deferred this to the prompts. This means that they cannot rule out, as far as I can see, the risk that the image generator itself is subject to this bias, where prompting it to generate more attractive faces will in fact make it generate more realistic faces. Without some control for this, I'm not sure that the paper achieves what it's set out to do sufficiently well and will possibly need to be complemented to be an authoritative resource for evaluating this bias.

    Read full reviewShow less
  2. This project is a solid effort in identifying the "Halo Effect" in models like PickScore and HPSv2. I really enjoyed the focus on beauty bias, though it does feel somewhat incremental given the existing literature on the subject. One thing I struggled with was the assumption that these reward models are accurate proxies for human liking; without validation against actual human preferences, it is hard to be certain. I also found it interesting, and a bit puzzling, that ImageReward showed nearly zero correlation while others were so high, and I would have loved to see a bit more exploration into why that happened.

    While I think framing this bias as a "manipulation" vulnerability feels a bit far-fetched in this specific context, I can see where the team was going with it. My main suggestion for the future would be to move away from the "synthetic loop" of using AI-generated images labelled by AI classifiers, as grounding these findings in real-world data is the only way to prove they aren't just artefacts of the generation process. However, I completely understand that this was a ton to cover within a hackathon timeline! Great job overall.

    Read full reviewShow less

Cite this project

@misc{mannering2026biased,
  title = {{Biased Attractiveness Bench (BAB): Image Reward Models Confuse Attractiveness with Realism}},
  author = {Harvey Mannering and Yaseen Mohammed Osman and Yeda WANG and Annika Catulli and Nehal Yasin},
  year = {2026},
  month = jan,
  note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/biased-attractiveness-bench-bab-image-reward-models-confuse-attractiveness-with-realism-sk7o}},
  url = {https://apartresearch.com/sprints/projects/biased-attractiveness-bench-bab-image-reward-models-confuse-attractiveness-with-realism-sk7o}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026