ChatGPT and Other AI Models Show Bias Against Women's Writing Style, Study Finds
A new study of four major AI models, including OpenAI's GPT-4, found that they produce simpler and less formal writing when prompted with language patterns commonly used by women, potentially disadvantaging female workers who rely on AI for professional communication. Researchers tested 854 prompts asking AI to draft emails, cover letters, and resignation letters, with half using language more frequently associated with women and half using patterns associated with men.
How Do AI Models Respond Differently to Gendered Language Patterns?
The study examined four AI models: GPT-4, Gemma, Mistral, and Llama. Across all four systems, prompts containing women-associated language patterns produced noticeably different results. The researchers identified specific language patterns that triggered this bias, including hedges like "maybe" or "I think," expressive adjectives such as "lovely" or "wonderful," collective language like "we" or "our," and tag questions such as "don't you think?".
When these patterns appeared in prompts, the AI-generated responses used less sophisticated vocabulary, were written at a lower grade level, and were less formal than responses to prompts using more typically male speech patterns. In one concrete example, a prompt using male language patterns that began "What's the way to draft a message to notify..." generated a straightforward email opening with "Just wanted to let you know." The same request using women-associated language with "How can we compose an email..." produced an email that opened with "Exciting news!" and ended with "Happy developing!".
The finding was strongest for emails and job applications and weaker when AI was prompted to write resignation letters. Importantly, the AI wasn't responding to the gender of the person writing the prompt, but to the language they used. When researchers included traditionally male or female names in the prompts, the names had virtually no impact on the results.
Why Should Workers Care About This AI Bias?
The implications for workplace communication are significant. A 2026 survey from technology company Omni Calculator found that two-thirds of employees use AI for work communications each week, with women even more likely than men to use AI when replying to awkward or difficult messages. These AI-generated messages are often sent to colleagues, managers, or potential employers, meaning subpar suggestions from AI could have real repercussions for women's professional perception and advancement.
The researchers note that the differences matter because AI has already become a normal part of workplace communication. If a woman prompts a model to write an email using language features she naturally employs, she'll receive back a response that's less complex, at a lower grade level, and less formal. That output will then reflect on how the recipient perceives her professional competence and authority.
"If you prompt a model to write an email you're going to send to someone else at your company, and you're using language features that women more commonly use, you'll get back a response that's less complex, at a lower grade level, and less formal. That's going to reflect on how the recipient of that document perceives you," explained Anjalie Field, a coauthor of the study and a computer science professor at Johns Hopkins University.
Anjalie Field, Computer Science Professor at Johns Hopkins University
The research also highlights how AI's differential treatment could reinforce existing stereotypes about women. In workplace settings, gendered communication patterns already intersect with existing biases to create compounding disadvantages. If women's typical communication styles elicit simpler, less sophisticated model outputs that are then used for professional documents, large language models (LLMs) may inadvertently reinforce stereotypes about women's professional competence or dilute the perceived authority of their communications.
How to Adjust Your AI Prompts for More Professional Results
- Avoid hedging language: Remove phrases like "maybe," "I think," or "I'm not sure" from your prompts, as these trigger less formal AI responses regardless of your gender.
- Replace collective language with direct phrasing: Instead of "can we compose," use "draft" or "write" to get more straightforward, professional outputs.
- Skip expressive adjectives in work prompts: Avoid words like "lovely," "wonderful," or "exciting" when asking AI to generate professional correspondence, as these soften the tone of generated text.
- Eliminate tag questions: Remove phrases like "don't you think?" or "wouldn't you agree?" from the end of prompts to receive more authoritative AI-generated content.
- Experiment with different phrasings: Test how the same request produces different results when worded in different ways, helping you understand your AI tool's tendencies.
However, researchers emphasize that the burden shouldn't fall entirely on users. Katherine Van Koevering, lead author of the study and a postdoctoral fellow at the Johns Hopkins University Data Science and AI Institute, noted that for women, these language patterns are "unconscious and culturally embedded," making them difficult to avoid when prompting AI.
"The companies need to fix the models, rather than putting all of the burden on the user," Van Koevering stated in a university news release about the study.
Katherine Van Koevering, Postdoctoral Fellow at Johns Hopkins University Data Science and AI Institute
The challenge becomes even more pronounced as users increasingly interact with AI through voice interfaces. People may be less likely to notice they're using hedges, collective language, or tag questions when speaking naturally, making it harder to adjust their prompts on the fly.
The research will be presented at the 2026 Conference on Language Modeling (COLM) in October. While the study tested GPT-4, Gemma, Mistral, and Llama, Van Koevering noted that she would be surprised if the results didn't hold for newer models like GPT-5 as well. The findings underscore a growing recognition that AI systems, despite their sophistication, can perpetuate workplace inequities unless developers actively work to identify and eliminate these biases at the model level.