A new study from Johns Hopkins University finds that several mainstream AI language models, including OpenAI's GPT-4 along with Meta's Llama, Google's Gemma, and Mistral, systematically make women's professional writing sound less intelligent than men's, even when the models never know who the writer is.
The research team asked the four bots to draft work emails, resignation letters, and job applications. For some prompts they used linguistic patterns more common among women, more expressive and inclusive phrasing, words like 'lovely,' 'we' instead of 'I.' For others they used masculine-coded language. The results were consistent across every model tested.
'If you prompt a model to write an email you're going to send to someone else at your company, and you're using language features that women more commonly use, you'll get back a response that's less complex, at a lower grade level, and less formal,' said senior author Anjalie Field, a Johns Hopkins computer scientist who studies ethics and discrimination in AI.
Crucially, the models were not simply mirroring the writer's tone. The quality gap persisted even after researchers controlled for it, suggesting the bots were actively reinforcing gendered assumptions rather than just reflecting their input. Traditional gendered names made no difference to the output; the linguistic cues alone were enough to trigger the double standard.
The findings add to a well-documented pattern of algorithmic bias, including prior studies showing AI models discriminate against speakers of African American Vernacular English. What makes this version of the problem especially insidious is its invisibility: an overtly sexist output is easy to spot, but subtly downgrading the sophistication of someone's professional writing is far harder to catch, particularly when users tend to assume AI outputs are smarter and more objective than their own.
As AI writing assistance becomes a default part of professional life, the study's implicit warning is that the technology does not just reflect the biases in its training data, it can quietly amplify them in the documents that shape careers.
Comments (0)
Log in to join the discussion
Log InNo comments yet