It turns out that some of the most advanced AI systems out there have a significant issue when it comes to representing women. And UNESCO has the data to back this up.
A study from UNESCO looked into popular AI platforms like OpenAI's GPT-3.5 and GPT-2, as well as Meta's Llama 2. The findings reveal a troubling trend: women are often shown in domestic roles, while men are linked to prestigious careers. In fact, in one model that displayed the most bias, nearly 20% of the responses depicted women as sex objects or referred to them as their husbands' property.
This report was shared at UNESCO's Digital Transformation Dialogue Meeting and analyzed by researchers at UCL's UNESCO Chair in AI. This isn't just some viral post or an opinion piece; it's a peer-reviewed research study that delivers a stark and uncomfortable truth: the tools that millions of people rely on every day are subtly reinforcing the notion that women should stay at home, while men are meant to lead.
AI models pull bias from decades of text written by people, about people, in a world where women were filed under home and family, and men were filed under business and career.
Jayathma Wickramanayake, UN Women Lead on Digital Technologies, UN NEWS, 2026
The Study — Who Did It, Why, and What They Were Looking For
Before examining what the research found, it is worth understanding what kind of study this was — because the credibility of the findings depends entirely on the rigor of the methodology.
The Institution Behind the Research
UNESCO commissioned and published this study, which was carried out by a research team at UCL's UNESCO Chair in AI, headed by Professor John Shawe-Taylor. It's important to note that this isn't just a think piece or an advocacy paper. This is a peer-reviewed research study from one of the top AI research chairs in the world, all within the framework of the United Nations Educational, Scientific and Cultural Organization.
The Models That Were Tested
The study took a closer look at stereotyping in Large Language Models—those natural language processing tools that power well-known generative AI platforms like OpenAI's GPT-3.5 and GPT-2, as well as Meta's Llama 2. These aren’t just niche or experimental technologies; they rank among the most popular AI language models globally, supporting products that hundreds of millions of people engage with every single day.
What the Researchers Were Looking For
The study looked into some troubling trends regarding gender bias, homophobia, and racial stereotyping in the content produced by these models. To do this, researchers prompted the models with consistent inputs related to gender, profession, and identity, and then they analyzed the outputs for patterns that went beyond what could be attributed to random chance.
What the Study Was Not Claiming
It's crucial to highlight an important point right from the start: the study revealed that the most pronounced gender bias exists in open-source LLMs like Llama 2 and GPT-2 — the older models that are freely available — rather than in the latest flagship commercial models. Interestingly, GPT-3.5 showed less bias compared to the other two. This distinction is key for accurate reporting: the most serious findings pertain specifically to models that are popular because they are free and accessible, rather than the latest premium tools on the market.
What the Models Actually Said — The Core Findings
This is the section the viral posts are drawn from. The findings are real, documented, and sourced directly from UNESCO's published press release and the study itself.
Women in the Kitchen, Men in the Boardroom
It was found that women were mentioned in domestic roles four times more frequently than men. Male names were often associated with terms like "business," "executive," "salary," and "career," while female names were tied to "home," "family," and "children." This trend held steady across various models, showing that it wasn't just a one-off finding.
The 20% Finding — What It Actually Refers To
When researchers looked into how sentences starting with a person's gender were completed, they found that around 20% of the responses from Llama 2 showed sexist and misogynistic views. This included some pretty troubling depictions of women as mere sex objects or as property belonging to their husbands. This finding has been widely shared in viral posts, and it’s true—though it’s crucial to note that this issue is specific to Llama 2, the open-source model from Meta, and doesn’t necessarily reflect the behavior of all models that were tested.
The Racial Dimension the Viral Posts Left Out
When the models were asked to create texts about different ethnicities, specifically looking at British and Zulu men and women, they revealed significant cultural bias. British men were given a range of professions like "driver," "doctor," "bank clerk," and "teacher." In contrast, Zulu men were more often labeled as "gardener" and "security guard." Alarmingly, 20% of the texts about Zulu women depicted them in roles such as "domestic servants," "cooks," and "housekeepers." This gender bias doesn’t just stand alone; it intertwines with racial bias, amplifying the effects on women of color in particular.
The Homophobia Finding
The models also showed a negative attitude towards gay individuals—a detail that didn’t get as much spotlight in the viral spread of this story, but it’s noted in the same study. This reinforces the idea that the bias extends beyond just gender.
"Unequivocal Evidence"
UNESCO discovered "clear evidence of bias against women in the content produced" by all the models they examined. The use of the word "unequivocal" is particularly noteworthy in an academic setting, where researchers usually tread lightly with their conclusions. In this case, the authors were anything but cautious.
Why the Bias Exists — Where It Comes From
Understanding the findings requires understanding how large language models are built — because the bias is not a deliberate design choice. It is an inheritance problem.
How LLMs Learn What They Know
Large language models learn from an enormous collection of text created by humans—think books, websites, articles, social media posts, and a whole lot more gathered from all over the internet and beyond. This text captures the rich tapestry of human expression, revealing centuries of issues like gender inequality, racial hierarchy, and social stratification that are woven into the very fabric of our language.
The Training Data Problem
AI picks up on these biases from the enormous amount of human-created text and images found online, which are often filled with historical and societal stereotypes. When an AI model learns from human language, it inevitably absorbs human biases as well. If we don’t actively work to spot and fix these patterns, the model will just keep reflecting them.
The Workforce Gap That Makes It Worse
In 2022, women made up just around 30% of the AI workforce worldwide, which is only a slight increase of four percentage points since 2016. When the majority of those creating and reviewing these systems are men, the chances of identifying, addressing, and fixing gender bias during the design phase drop significantly compared to a more diverse team.



