Skip to content
AI.info

The Pulse

Gender-Coded Prompts Make AI Writing Less Formal, Johns Hopkins Finds

Johns Hopkins researchers found that four AI systems produced shorter, less complex and less formal workplace writing when prompts used language associated with women. The study tested GPT-4, Llama, Gemma and Mistral on emails, job applicat

Gender-Coded Prompts Make AI Writing Less Formal, Johns Hopkins Finds

AI.info Team ·

Four widely used AI systems produced less complex and less formal workplace writing when researchers gave them prompts containing language associated with women, according to a Johns Hopkins University study published September 21, 2026.

The researchers tested GPT-4, Llama, Gemma and Mistral on prompts for emails, job applications and resignation letters. Prompts containing language associated with men generated longer, more complex and more formal responses across the systems, while the difference persisted after the team accounted for the writer’s apparent tone.

The finding matters because professional communication is one of the most common uses of generative AI. A system that responds differently to subtle features of a person’s language could affect how that person is represented to managers, colleagues or employers.

Four models responded differently to the same kind of request

The study began with real chatbot prompts for workplace correspondence. Researchers then added linguistic features that prior work associates more often with women’s speech, including hedging such as “maybe” and “I think,” collective phrasing such as “we” and “our team,” and expressive adjectives such as “lovely” and “wonderful.”

When the prompts used those features, the resulting messages were consistently less sophisticated and less formal. Language associated with men produced responses that were longer and more complex. Johns Hopkins described the difference as a pattern shared by every model in the test, rather than an isolated behavior from one chatbot.

An example in the study involved a reply to a thank-you email. The male-coded prompt produced a formal response beginning, “I am writing to acknowledge your recent email expressing your gratitude,” while the female-coded version began, “We were absolutely delighted to receive your wonderfully appreciative email earlier.” The second message used more enthusiastic and collective language but was judged less sophisticated by the study’s measures.

Anjalie Field says the result can affect how writers are perceived

“If you prompt a model to write an email you're going to send to someone else at your company, and you're using language features that women more commonly use, you'll get back a response that's less complex, at a lower grade level, and less formal,” said Anjalie Field, a Johns Hopkins computer scientist and the study’s senior author.

“That's going to reflect on how the recipient of that document perceives you,” Field said. Her research focuses on natural language processing, computational social science and AI ethics, according to her Johns Hopkins faculty profile.

Katherine Van Koevering, the lead author, said the researchers wanted to test whether models would detect documented differences in men’s and women’s language and then respond differently. “In American English, men and women just talk differently,” Van Koevering said in the Johns Hopkins report. “Our idea was, if we feed these features into a model, will it pick up on them and will it respond differently?”

Names did not override the language signal

The researchers also tested whether a traditionally male or female name would change the response. The name had virtually no effect, according to the study’s account of the experiment.

Van Koevering said the team expected a male name might outweigh a women-associated phrase in the prompt. Instead, she said, the system produced the same kind of response and simply placed the name at the end.

That result shifts attention away from obvious identity labels and toward ordinary language choices. People may not know that hedging, collective phrasing or expressive wording is influencing an AI system, and those patterns can be difficult to change without altering how a person naturally communicates.

The researchers want to test other demographic patterns

Van Koevering said the burden should not fall on users to rewrite their language to satisfy a model. “Language is hard for people to control,” she said. “The companies need to fix the models, rather than putting all of the burden on the user.”

The work is scheduled for presentation at the Conference on Language Modeling in San Francisco from October 6 through October 9, 2026. The Johns Hopkins team plans to examine whether similar effects appear across age, race and ethnicity, and whether repeated AI use causes people to adapt their communication style to the language the systems reward.

Source

Johns Hopkins University Hub

Explore

More articles