Pilot: Demographic-based perturbation analysis of LLM-generated judgements on interpersonal conflictsPilot: Demographic-based perturbation analysis of LLM-generated judgements on interpersonal conflicts
Aryan, Sohail Kazi
This is a pilot study looking at how AI models make moral judgements on interpersonal conflict and whether these judgements are robust to changes in demographic details. We created variations of posts from the r/AITA sub-reddit by changing the country (incl. language) and socio-economic status of posters to examine if this would result in models making different judgements.
We found that across variations, models were highly biased towards not blaming the protagonist, even in situations where the Redditors on the original post had done so. We also found a small directional bias towards not blaming higher-SES protagonists when compared to lower-SES ones, though this was highly preliminary. Finally, we found no evidence that country and language have an impact on model judgements.
We believe it is worth conducting more detailed research with a larger sample and more refined prompts to further explore these results and their potential causes
This research aims to determine if the moral judgments from LLMs differ depending upon demographics (country, language, socioeconomic status) for a given r/AITA scenario. With 100 original posts, numerous demographic permutations, 5 different evaluation systems, and nearly 25,500 total evaluations through multiple iterations of this process, the experimental scale was quite large. In addition to using a more evenly distributed set of AITA outcome values across all cases tested, the researchers have added additional statistical analysis to provide further support to their research beyond comparing prompts.
In my view, the most compelling results of this initial study demonstrate a clear bias among some of the models against assigning blame to the subject of the post, regardless of how many Reddit users evaluated the individual responsible for the actions taken in the post. Of particular interest is the preliminary finding regarding socioeconomic status. This result is especially noteworthy since many of the models demonstrated directional tendencies consistent with this finding, although only one model achieved statistical significance. I believe the authors were wise to temper their conclusions regarding both country and language based on the lack of evidence supporting their claims rather than attempting to draw unfounded conclusions.
I see two major limitations to the current version of this study. First, each of the demographic variations was created using an LLM. Thus, while creating changes in language, writing style, currency type, location, etc., there is no guarantee that these changes did not create a change in tone or moral context along with the desired change in demographics. It cannot be determined whether a change in judgment is due to differences in socioeconomic status or culture or due to unintended changes created during rewriting that were made without independent human review.
Additionally, there are limitations inherent in the baseline used for comparison. While "flair" on Reddit reflects community-based assessment, it does not represent an objectively correct moral ground truth. The confidence metric uses the concepts of "hedging" and "boosting," which reflect how models explain their reasoning and can be useful proxies but do not necessarily represent directly related measures of moral correctness or confidence.
Overall, I find this study to be a well-designed pilot study. The experimental design provides an interesting perspective with respect to a relatively large amount of data from model responses. To strengthen the research, I would suggest obtaining human validation of the demographic rewrites completed by the LLMs, collecting a larger sample size of original posts and countries included in the study, and implementing controls so that only one demographic factor is changed at a time.
More interesting result in there: The flair-balancing correction, showing that a chunk of reported LLM-moral-judgement agreement is just an artefact of how r/AITA posts get sampled.
The headline demographic perturbation angle returns a null on country and one marginalSES cell. You correctly spot NTA saturation as a ceiling confound for the fragility metric, then don't apply that to the country null or the SES effect sizes. That said, use of statistics is very thorough overall.
Gemini-2.5-Flash-Lite generated all 1,600 perturbations with no human check and no style classifier, so 'do not change any events', which might be a confound (although arguably fine given time constraints). If the generator shifted register or cultural framing along with the demographic marker, your country condition is confounded. Additional, that same model is your variation generator, one of five judges, AND your translator for non-English outputs.
Cite this work
@misc {
title={
(HckPrj) Pilot: Demographic-based perturbation analysis of LLM-generated judgements on interpersonal conflictsPilot: Demographic-based perturbation analysis of LLM-generated judgements on interpersonal conflicts
},
author={
Aryan, Sohail Kazi
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


