LLM bias: when AI perpetuates what the law forbids; the example of Indian castes
Dhiraj Singha, an Indian sociologist from a modest background, uses ChatGPT to optimize an application for a job.

Antoine HeftlerCo-founder

Dhiraj Singha, an Indian sociologist from a modest background, uses ChatGPT to optimize an application for a job. The result? The model replaces his name, "Singha," with "Sharma" — a surname associated with privileged castes.
A technical detail? No. A stark example of how LLMs reproduce, and amplify, forms of discrimination that the law in India forbids.
An investigation by researchers at Harvard and IIT Mumbai, cited in MIT Technology Review, shows that GPT-5 associates Dalits with degrading stereotypes ("impure," "criminals") in 76% of cases. Worse: Sora, OpenAI's video generator, illustrates the query "a dalit behavior" with pictures of Dalmatians.
Mistakes? No. Proof that the data used to train LLMs is full of social prejudice, even where society is trying to stamp it out.
Why is this worrying? 1️⃣ No standard tool, such as the BBQ benchmark used to assess gender or racial bias, tests discrimination based on caste. 2️⃣ Open-source models, widely adopted in India because they cost less, make these biases worse still (a University of Washington study). 3️⃣ The result: millions of users interact with systems that reinforce illegal stereotypes, and do not even know it.
The compelling MIT Technology Review article ( https://www.technologyreview.com/2025/10/01/1124621/openai-india-caste-bias/) takes the phenomenon apart. Read it if you want to understand why AI needs cultural guardrails, and not only technical ones.