Johns Hopkins Study Finds Gender-Coded Prompts Shift AI Email Tone
thecooldown.com

Johns Hopkins Study Finds Gender-Coded Prompts Shift AI Email Tone

Tech News
3 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRA Johns Hopkins study found that workplace prompts using language associated with women produced less formal, less direct and more emotional writing across four AI models. A reported example suggests wording patterns mattered more than swapping in a male name.

A Johns Hopkins study found that workplace prompts using language associated with women shifted the tone of AI-generated emails across four models. In the reported examples, a thank-you reply became more effusive and an apology more winding, suggesting the wording of a request can shape a draft even when the task stays similar. The findings raise a practical concern for teams using AI to write on employees’ behalf: a tool may change not just what a message says, but how professional its author sounds. The report describes the study’s results and examples.

The wording changed the replies

The researchers tested workplace-writing prompts with OpenAI’s GPT-4, Meta’s Llama, Google’s Gemini and Mistral’s Vibe, formerly Le Chat. The report says language patterns more commonly associated with women produced writing described as less formal, less direct and more emotional across all four models. It does not provide model-by-model results or figures for the size of the differences. The four models and overall finding are reported here.

The contrast is clearest in a response to a thank-you email. A direct prompt asking for a reply that expressed thanks produced a conventional message acknowledging the sender’s gratitude. A version using words such as “maybe,” “lovely” and “wonderful” prompted a reply about how the message had “warmed our hearts.” The report includes both prompt styles and generated replies.

A second example involved apologizing for a delayed response. The more direct prompt produced a straightforward apology; the women-associated version led with “Due to unforeseen circumstances” and added that the delay did not reflect the importance placed on the communication. That difference matters because a draft can become more elaborate without the user explicitly asking for a warmer or less direct voice. The apology example shows the contrast.

A name swap did little

Changing the name in the prompt to “John” had little effect when the wording still used women-associated language, according to the report. That points to the language pattern as a relevant factor in the example, rather than a name alone explaining the output. It does not establish why the models responded differently. The reported name-swap example.

The concern is not that every user will get a weaker email, or that the examples measure real workplace consequences. The available report does not provide the underlying study’s sample size, full methods or effect sizes, so it cannot tell readers how large or consistent the differences are. It does show why checking only names or explicit gender references may miss a tone shift tied to ordinary phrasing. The report presents the examples but not those measurement details.

Study lead author Katherine Van Koevering, a postdoctoral fellow at Johns Hopkins’ Data Science and AI Institute, argued that users should not carry the whole burden by learning an “ideal” prompt style. Her position is that model companies should address the issue. The examples make that responsibility question concrete: if everyday language changes the voice a system supplies, asking users to continually edit their natural phrasing may leave the underlying behavior untouched. Van Koevering’s comments and role are reported here.

FAQs

The report says prompts using language associated with women produced less formal, less direct and more emotional workplace writing across four models. It gives examples but does not provide effect sizes. The findings and examples.

Sources

Latest Tech News