Skip to content
AI.info

Research

Exploring the Interplay Between Voice, Personality, and Gender in Human-Agent Interactions

Overview Research area: Human-Computer Interaction / Human-Agent Interaction, specifically voice-based artificial agents and personality perception. Technical level: Intermediate. The paper combines s

arXiv
2602.10535
Published
2026-02-11
Authors
Kai Alexander Hackney, Lucas Guarenti Zangari, Jhonathan Sora-Cardenas, Emmanuel Munoz, Sterling R. Kalogeras, Betsy DiSalvo, Pedro Guillermo Feijoo-Garcia

AI summary

Overview

Research area: Human-Computer Interaction / Human-Agent Interaction, specifically voice-based artificial agents and personality perception.

Technical level: Intermediate. The paper combines study design, psychometric instruments (TIPI), and non-parametric statistics (Mann-Whitney U, Spearman correlation, Holm-Bonferroni correction), but the framing is conceptual rather than engineering-heavy.

Scope: An exploratory user study with 388 participants who evaluated four synthetic voices (male/female x introverted/extroverted) to test whether voice-only agents convey extroversion and whether users project their own personality onto agents.

What This Paper Is About

Designers of voice-based agents want to know whether personality can be conveyed through voice alone, with no visual embodiment, and whether users respond better to agents whose personality resembles their own. Prior work on personality alignment has mostly used limited vocal features or combined voice with visual and behavioral cues, leaving voice-only perception poorly understood. This paper tests whether users can distinguish introverted from extroverted synthetic voices, and whether user-agent personality synchrony shows up in those judgments.

Key Contributions

  1. An exploratory study with 388 participants comparing perceived extroversion across four synthesized voices (male introverted, male extroverted, female introverted, female extroverted) generated from human recordings, each evaluated in a voice-only modality.
  2. Evidence that participants differentiated extroversion across the female voice conditions but not across the male voice conditions, an asymmetry observed consistently across the analyses in the study.
  3. Preliminary evidence consistent with personality synchrony during the first agent interaction, with a significant correlation among male participants that survived Holm-Bonferroni correction, plus an observed anchoring pattern where second-agent evaluations tracked first-agent perceptions rather than self-reported personality.
  4. Methodological insights and explicit caveats about stimulus diversity, arguing that one synthesized voice per condition cannot separate intended personality cues from idiosyncratic vocal properties such as accent, speaking style, or other sociophonetic characteristics.

Main Findings

  • Female voices were differentiated: Participants evaluating female voices (n=195) showed a significant negative Spearman correlation between perceived extroversion and intended personality condition (r_s = -0.533, n = 195, p < .001).
  • Male voices were not differentiated: Participants evaluating male voices (n=193) showed no significant association (r_s = 0.0083, n = 193, p = 0.30).
  • First-exposure comparison confirmed the female effect: Participants who heard the introverted female voice first (n=95) differed significantly from those who heard the extroverted female voice first (n=100) in perceived extroversion (Mann-Whitney U = 7915, p < .001).
  • First-exposure comparison found no male effect: Participants who heard the introverted male voice first (n=94) did not differ significantly from those who heard the extroverted male voice first (n=99) (Mann-Whitney U = 4576, p = 0.85).
  • A possible synchrony effect in the full sample did not survive correction: Across all 388 participants, self-reported extroversion correlated positively with perceived extroversion of the first voice (r_s = 0.124, n = 388, p = 0.025), but this did not remain significant after Holm-Bonferroni correction (adjusted alpha = 0.017).
  • Synchrony was significant among male participants: Male participants showed a correlation between their own and the first agent's perceived extroversion (r_s = 0.150, n = 276, p = 0.013) that remained significant after correction.
  • No synchrony among female participants: Female participants showed no significant correlation (r_s = 0.025, n = 102, p = 0.80), though the authors note the female subgroup was substantially smaller than the male subgroup.
  • Synchrony appeared only at first interaction: For the second agent evaluated, no significant relationship emerged between self-reported extroversion and perceived agent extroversion. Instead, a significant relationship appeared between the perceived extroversion of the first and second agents, which the authors describe as consistent with an anchoring effect.
  • A possible confound in the male stimuli: The introverted male voice exhibited a noticeable Southern U.S. accent, which may have influenced perceptions independently of extroversion.

Methodology in Plain English

The study ran in Spring 2025 at a large Southeastern institution in the United States and was approved by the Georgia Institute of Technology IRB (Protocol No. IRB2025-303).

A preliminary data collection recruited 346 adults through Prolific, restricted to English-fluent participants located in the United States. Personality and demographics were collected, and participants could optionally submit voice recordings reading standardized passages. Thirteen participants submitted recordings via the Phonic platform; two were excluded for failing audio quality and script-adherence criteria, leaving 11 usable recordings. Ages ranged from 20 to 61 years, with 6 male and 5 female participants; one participant also reported frequent use of Hindi.

From those 11 recordings, the researchers selected one participant per condition using the highest and lowest TIPI extroversion scores within each gender. For male participants, one individual scored the minimum (1) and one the maximum (7). For female participants, one individual scored the maximum (7) and the lowest available score was 2.35. Each selected recording was manually reviewed for audio quality and script adherence.

Synthetic voices were then generated with ElevenLabs using a single model (Eleven Multilingual V2) and identical parameter settings (Stability: 50%, Similarity: 75%, Style Exaggeration: 0%, Speaker boost: on). All voices read the same script, adapted from Job Burnout, chosen as a continuous standardized narrative; it lasted an average of four minutes and 42 seconds across conditions.

The main study was asynchronous and fully digital via Qualtrics. Of 429 recruited adults, 388 completed the survey and provided informed consent. Participants came from a single course in the College of Computing at the Georgia Institute of Technology. Each participant was randomly assigned to one of four groups determining voice gender and presentation order: extroverted male then introverted male; introverted male then extroverted male; extroverted female then introverted female; or introverted female then extroverted female. Every participant listened to two voices of the same gender differing in extroversion, hearing each recording in full before completing the questionnaire. The estimated completion time was 20 minutes.

Extroversion was measured with the Ten Item Personality Measure (TIPI) for both self-assessment and perceived agent personality. Because the scores come from Likert-type responses and were treated as ordinal, the analysis used non-parametric methods: Mann-Whitney U tests for between-group and paired comparisons, and Spearman's correlation for associations. The Holm-Bonferroni method was applied to address multiple comparisons, with a significance threshold of alpha = 0.05.

Why This Matters

The paper shows that the same synthesis pipeline can produce voices that are perceived very differently depending on gender condition, which complicates the assumption that a personality label attached to a voice will reliably reach the user. It also frames synchrony as something that may be strongest at first contact and then give way to anchoring on earlier impressions, a finding with direct design implications for how many agents a user encounters and in what order.

Real-world applications:

  • Voice assistants and customer service agents, where designers choose whether an agent should sound introverted or extroverted and whether that choice will be perceived at all.
  • Mental health and wellness agents, a domain the paper's background section repeatedly references as one where user-agent and user-designer similarity shape trust and engagement.
  • Educational and tutoring agents, which the authors name as a context for future synchrony research.
  • Speech interfaces in vehicles, building on cited prior work showing that drivers trusted and liked assistants whose personality matched their own.

Industry relevance: the findings argue against relying on a single voice sample to represent a personality type, and for controlling or validating acoustic features such as pitch, tempo, and prosody before shipping gendered personality cues. The anchoring result also implies that a user's first agent interaction may color every subsequent one, which matters for products that expose users to multiple agent personas.

Future Directions

  • Increase the number and diversity of voice stimuli, with multiple synthesized voices per experimental condition, to disentangle personality expression from idiosyncratic vocal characteristics.
  • Objectively validate acoustic features such as pitch, tempo, and prosody to confirm that intended personality manipulations are actually present in the audio.
  • Investigate when and how user-agent personality synchrony emerges, including its dependence on user traits, interaction context, and how much information the user has about the agent.
  • Examine the observed gender asymmetry and the male-participant synchrony result with larger, more balanced participant subgroups, since the current female subgroup (n=102) was much smaller than the male subgroup (n=276).
  • Explore these dynamics in applied settings such as education and mental health support.

Target Audience

Researchers and practitioners in human-agent interaction, voice interface design, and conversational AI who need to understand the limits of conveying personality through synthetic speech. Designers specifying agent voice personas will benefit from the cautionary framing about stimulus diversity and the anchoring behavior. Students and newcomers to HCI study design will find a clear example of an exploratory study that reports its own confounds candidly.

Authors’ abstract

To foster effective human-agent interactions, designers must understand how vocal cues influence the perception of agent personality and the role of user-agent alignment in shaping these perceptions. In this work, we examine whether users can perceive extroversion in voice-only artificial agents and how perceived personality relates to user-agent synchrony. We conducted a study with 388 participants, who evaluated four synthetic voices derived from human recordings, varying by gender (male, female) and personality expression (introverted, extroverted). Our results show that participants were able to differentiate perceived extroversion in female agent voices, but not consistently in male voices. We also observed evidence of perceived personality synchrony, particularly in participants' evaluations of the first agent encountered, with this effect more pronounced among male participants and toward male agents. We discuss these findings in light of limitations in stimulus diversity and voice representation, and outline implications for the design of voice-based agents, particularly regarding the interaction between gender, personality perception, and initial user impressions. This paper contributes findings and insights to consider the interplay of user-agent personality and gender synchrony in the design of human-agent interactions.

Read the original paper