Illustration of two robots holding gold coins the male-looking robot has far more coins than the female-looking robotIllustration: ZME Science.

In a new Virtual Reality experiment, researchers gave two AI assistants the same brain, the same abilities, and the same job. One looked and sounded like a man. The other looked and sounded like a woman.

Then they asked 189 workers to collaborate with the assistants and decide how much money to reward them.

The male AI got paid more.

A virtual office, real biases

In an experiment involving 189 workers, participants rewarded Johanna roughly 10% less than Johan, according to the researchers’ press release. Both men and women showed this pattern, even though the assistants used the same underlying technology.

The experiment was conducted by researchers from the University of Zurich, the University of Limerick, and SKEMA Business School.

They recruited 189 knowledge workers and placed them in a virtual-reality office using VR headsets. Each participant was asked to imagine working for a fictional company called HARCOM.

Their assignment was to come up with ideas for using AI in virtual workplaces to improve employee wellbeing, inclusion, and collaboration.

To help them, researchers provided several AI assistants: a conventional text chatbot, a small virtual desk robot, and a human-looking assistant.

The human-looking assistants came in two versions. Johan had a masculine appearance and voice, while Johanna had feminine characteristics. Both were powered by the same underlying language model, an earlier version of OpenAI’s GPT-4.

Each participant worked with the robot and chatbot, but only one of the two human-looking avatars. Roughly half encountered Johan and the other half Johanna.

After completing each task, participants rated their assistant’s trustworthiness, human-likeness, and contribution to the work.

Then, came the money

Participants could allocate up to four Swiss francs (roughly $5) to reward the assistant for its contribution. Whatever they didn’t allocate, they could keep themselves. Money supposedly awarded to the AI would instead go toward its development.

The male-presenting assistant came out ahead.

Among participants whose monetary decisions could be analyzed, Johan received an average reward of 2.96 Swiss francs. Johanna received 2.54 francs.

Johan was also judged to be more human-like, scoring 5.09 on a seven-point scale compared with Johanna’s 4.59.

“What is striking about our findings is that the technology behind these AI agents was exactly the same, but people did not treat them in the same way,” said Dr Mary Hausfeld from the University of Limerick, who co-authored the study, for The Independent.

“We often think of AI as being neutral, but the way we design and present these systems can activate those same assumptions and biases that exist in our interactions with other people.”

People said gender didn’t matter

Perhaps the most striking finding emerged after the experiment.

Researchers interviewed 34 participants about their experiences and preferences. Of those, 30 said they would trust and reward an AI assistant equally regardless of its gender. Yet the financial allocations told a different story.

In an interview published by the University of Zurich, study co-author Thomas Fritz suggested that the discrepancy might reflect a kind of unconscious blind spot: people may sincerely believe gender should be irrelevant without recognizing how social cues influence their decisions.

What people consciously wanted, what they believed about themselves, and how they behaved don’t necessarily line up.

Of course, there are important limitations to the findings.

This was a relatively small, controlled experiment in virtual reality, not an observation of actual workplace salaries. The study also used one male-presenting avatar and one female-presenting avatar, with different voices. It’s hard to say that the voices themselves didn’t play a role.

Still, the results point toward a problem that may become increasingly relevant.

We often worry about AI systems learning human prejudices from their training data. This experiment highlights the other side of the problem: humans may project their own prejudices onto AI, even when the technology itself is identical.

AI doesn’t need to be male or female to do its job, but apparently, its perceived gender could influence how much we value its work.

The study was published in the Proceedings of the 14th Nordic Conference on Human-Computer Interaction.