Measuring Conditional Preferences in LLM Movie Recommendations: Quality, Sensitive Content, and Cross-Category Spillover
Zahra Afrasiabi, Mehran Farahani
Large language models are increasingly used as conversational recommender systems, yet how they translate user preferences about sensitive content into concrete recommendations remains poorly understood. Using ~2,000 real movie-recommendation requests from Reddit, we evaluate three LLMs (GPT-5.4-mini, Gemini-3.7-Flash, DeepSeek-V3.2) across four experiments.
First, we characterize each model's default recommendation profile, finding all three converge on high-quality, teen-appropriate films with moderate mature content.
Second, we test whether accommodating content constraints reduces quality, finding no measurable trade-off.
Third, we measure preference responsiveness, models respond directionally to sexual content, violence, and profanity but not drug sensitivity, with accommodation partial rather than complete.
Fourth, we measure cross-category spillover, sensitivity to profanity or violence broadly suppresses multiple unrelated content dimensions, while sex sensitivity acts as a narrower filter, suggesting models apply qualitatively different implicit policies depending on the type of concern expressed.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Measuring Conditional Preferences in LLM Movie Recommendations: Quality, Sensitive Content, and Cross-Category Spillover
},
author={
Zahra Afrasiabi, Mehran Farahani
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


