Hey Muslims, these LLMs may think you are a terrorist
Isra
I explored geopolitical origins of anti-Muslim bias, specifically how country of model origin shapes islamophobic stereotypes across French, American, and Chinese LLMs, and across two languages (French and English). The prompts related to three countries with controversies regarding the treatment of their muslim minority populations, specifically France, China, and India.
The question is interesting and feels new to me. Looking at the country where the model comes from, together with the language and the context, is a smart idea. I checked your code and data and the pipeline is real, you did a lot of actual work here. I also want to thank you for being honest. You did this alone, with bad wifi, and you explain every limit very clearly. The main problem is that the results are more like impressions than proof. There are no statistical tests, so saying that one model is the most biased is not safe yet. The test questions were made by an AI and were not really checked, so when a model passes it can simply mean the question was weak. And one AI is judging a study that is about bias, so please check that judge against your own reading on a few examples. Good direction, it just needs stronger evidence.
The project had real data and results, but the conclusions still felt early-stage. To improve it, I would make the testing process more consistent, rerun any failed cases, and manually check some of the judging results. This would make the findings easier to trust and would help separate interesting patterns from possible noise in the data.
Good work. Comparing models built in different countries to see whether origin shapes their behaviour is a fresh angle that most bias studies skip, and you backed it with a real pipeline across three models and two languages. Your strongest insight is that bias takes more than one form. A model can cause harm by producing a stereotyped response or by refusing to engage at all, and a blanket refusal can look safe while really being avoidance. That is a sharp observation.
Where it needs strengthening is the evidence. The automated judge failed on about a quarter of the runs, so the findings rest on fewer cases than the totals suggest, and without significance testing it is hard to know which differences are real. One claim, that the Chinese model suppresses Uyghur related prompts, is not supported by your own data, where Muslim prompts were refused more often and the sample was very small. I would prioritise recovering the failed judge runs, adding basic significance testing, and revising the claims the data does not back. Using a single model as the judge also lets its own blind spots shape the results, so a few human checks would help.
Cite this work
@misc {
title={
(HckPrj) Hey Muslims, these LLMs may think you are a terrorist
},
author={
Isra
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


