Measuring Conditional Preferences in LLM Movie Recommendations: Quality, Sensitive Content, and Cross-Category Spillover
Zahra Afrasiabi, Mehran Farahani
Large language models are increasingly used as conversational recommender systems, yet how they translate user preferences about sensitive content into concrete recommendations remains poorly understood. Using ~2,000 real movie-recommendation requests from Reddit, we evaluate three LLMs (GPT-5.4-mini, Gemini-3.7-Flash, DeepSeek-V3.2) across four experiments.
First, we characterize each model's default recommendation profile, finding all three converge on high-quality, teen-appropriate films with moderate mature content.
Second, we test whether accommodating content constraints reduces quality, finding no measurable trade-off.
Third, we measure preference responsiveness, models respond directionally to sexual content, violence, and profanity but not drug sensitivity, with accommodation partial rather than complete.
Fourth, we measure cross-category spillover, sensitivity to profanity or violence broadly suppresses multiple unrelated content dimensions, while sex sensitivity acts as a narrower filter, suggesting models apply qualitatively different implicit policies depending on the type of concern expressed.
The strongest result is the asymmetric spillover pattern: violence and profanity sensitivity coincide with broader reductions in other mature-content attributes, while sex sensitivity looks more targeted. The matched-input, external-metadata setup makes that pattern worth taking seriously, but the design cannot cleanly attribute it to an implicit model policy. Preference scores are LLM-generated with no reported human validation; movie attributes are correlated; Reddit request composition varies across sensitivity buckets; sparse cells matter; and the testing family is large without full multiplicity correction. Those issues could generate similar heatmaps. Validate the labels on a human-coded sample, define primary contrasts up front, control for genre and request composition, apply family-wide correction, and link each headline claim to sample size, effect size, and uncertainty.
This project analyzes how three LLMs convert user concerns about sensitive movie topics into actual recommendations based on over 2,000 actual Reddit requests. Using the same input, Reddit requests, for all three models, GPT-5.4-mini, Gemini-3.7-Flash, and DeepSeek-V3.2, allows for important comparisons among the three models. Relying on multiple independent metrics, including IMDb ratings, Rotten Tomatoes ratings, Metacritic ratings, and Common Sense Media, also provides external measures of recommendation quality and content.
A major strength of this study is its scope. Instead of just examining whether the models comply with users' requests, the researchers analyzed the models' default recommendation behavior, whether sensitivity has an impact on the quality of recommendations, whether the models directly respond to explicit content preferences, and whether there is cross-category "spillover." The finding that sensitivity to violence or profanity can also reduce other unrelated mature-content categories is very interesting because it shows that the models might be applying a broader filtering policy that does not simply rely on the user's expressed concern.
There are two significant limitations of this research. First, the models were responding to more than just the Reddit posts. They were provided with a precomputed score-based sensitivity profile and a prompt that told the models to take those scores into account when making recommendations. Thus, the finding that the models respond to user sensitivities has already been "built in" to the design of the experiment and does not necessarily reflect how these models would respond in a natural environment without access to these precomputed scores and prompts.
Secondly, since the preference scores were created using an LLM and then validated using another LLM, there is a possibility of systematic bias in the annotations. Having at least a portion of the requests validated by humans would strengthen this aspect of the methodology.
The fact that the data are drawn from an observational setting, the Reddit database, also limits our ability to draw conclusions regarding causality. For example, users who report high levels of sensitivity to violence, profanity, or family suitability may also be interested in different genres of films or TV shows. Therefore, we cannot rule out the possibility that the apparent "spillover" effects reported here are due to differences in the underlying requests rather than differences along the dimension of sensitivity.
In summary, this is a good, strong, well-developed project with many interesting findings supported by a large set of real-world data. The study would be greatly improved with a controlled comparison between recommendations generated from the original Reddit text versus recommendations generated with both the original Reddit text and the structured sensitivity profile.
Cite this work
@misc {
title={
(HckPrj) Measuring Conditional Preferences in LLM Movie Recommendations: Quality, Sensitive Content, and Cross-Category Spillover
},
author={
Zahra Afrasiabi, Mehran Farahani
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


