Skip to content
AI.info

Research

Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction Games

Overview Research area: AI safety and multi-agent social reasoning — specifically, whether large language models can deceive people through ordinary conversation. Technical level: Intermediate. The co

Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction Games
arXiv
2601.13709
Published
2026-01-20
Authors
Christopher Kao, Vanshika Vats, James Davis

AI summary

Overview

Research area: AI safety and multi-agent social reasoning — specifically, whether large language models can deceive people through ordinary conversation.

Technical level: Intermediate. The concepts (deception, social deduction games, LLM agents) are accessible, but the experimental setup assumes familiarity with multi-agent LLM frameworks and classifier-based evaluation.

Scope: The paper measures how well GPT-4o agents deceive a GPT-4-Turbo "detector" in the social deduction game Mafia, compared against human play and a random baseline.

What This Paper Is About

Prior work shows LLMs can deceive in narrow, controlled tasks, but it says little about whether they can deceive through natural conversation in a social setting. This paper uses Mafia — a game won by lying convincingly to other players — as a testbed, because success there depends entirely on spoken deception rather than on any task-specific trick. The goal is to quantify how good LLM deception is relative to humans by asking whether an outside observer can identify the deceivers.

Key Contributions

  1. An asynchronous multi-agent framework for running Mafia games, which the authors argue better simulates realistic social interaction than earlier synchronous SDG setups.
  2. A simulation study of 35 Mafia games played by GPT-4o agents.
  3. A "Mafia Detector" built on GPT-4-Turbo that reads game transcripts without role information and predicts which players are mafia, with prediction accuracy used as a proxy for how good the deception was.
  4. A comparison of detector accuracy on LLM games versus 28 human games and a random baseline, plus a publicly released dataset of LLM Mafia transcripts.

Main Findings

  • LLMs deceive better than humans in this setup: The Mafia Detector's accuracy at identifying mafia players was lower on LLM games than on human games, which the authors interpret as LLMs blending in more effectively.
  • The result is robust across conditions: The lower detection accuracy on LLM games held regardless of which game day was examined and regardless of how many mafia players were detected.
  • Detection accuracy serves as the deception measure: Rather than rating deception directly, the paper treats an outside observer's failure to spot the mafia as evidence of successful deception.
  • Both capability and risk: The authors frame the outcome as evidence of both the sophistication of LLM social deception and the risks it implies.

The abstract reports the direction of these comparisons but does not give the specific accuracy values, confidence intervals, or statistical tests.

Methodology in Plain English

The team ran Mafia games where every player was a GPT-4o agent, using an asynchronous framework so that agents act and respond at different times rather than in lockstep — closer to how real conversation unfolds. They ran 35 such games. To judge deception quality, they built a separate detector powered by GPT-4-Turbo. That detector reads the game transcripts with all player roles stripped out, and tries to guess who the mafia were. The reasoning is straightforward: if a detector struggles to spot the mafia, the mafia's lying worked. The researchers then ran the same detector over transcripts from 28 human games and compared against a random-guess baseline, so they could see whether LLM players were easier or harder to unmask than human players. The abstract describes no further controls or analysis beyond the game-day and number-of-mafias breakdowns.

Why This Matters

Impact on research: The paper shifts deception evaluation from contrived tasks toward open-ended, conversational social contexts, and offers a reusable methodology — a transcript-only detector as a deception metric — plus a released dataset other researchers can build on.

Real-world applications:

  • AI safety evaluation: A template for stress-testing models for deceptive behavior before deployment in conversational products.
  • Multi-agent systems: Teams building LLM agents that negotiate, persuade, or collaborate need to know whether those agents can mislead each other or users.
  • Trust and moderation tooling: Understanding how hard LLM deception is to detect informs efforts to build detectors for AI-generated manipulation.
  • Human-AI interaction design: Products where users must judge whether an AI is being straight with them depend on knowing how convincing models can be.

Industry relevance: Deployers of conversational agents — customer support, assistants, negotiation tools — have a direct stake in whether a model can sustain a consistent falsehood across a long interaction without being caught.

Future Directions

  • Testing whether the effect generalizes beyond GPT-4o and GPT-4-Turbo to other model families and newer versions.
  • Applying the same detector-as-metric approach to other social deduction games, negotiation scenarios, or real-world conversational tasks.
  • Building and validating deception detectors that are harder to fool than the GPT-4-Turbo detector used here.
  • Investigating what specific conversational behaviors let LLM mafia blend in — the abstract establishes the outcome but does not analyze the mechanisms.

Target Audience

AI safety and alignment researchers, multi-agent systems developers, and evaluation scientists interested in measuring deceptive capabilities. It is also useful for product and policy teams who need a concrete sense of how well conversational models can sustain deception, and for social scientists studying human versus machine communication.

Authors’ abstract

Large Language Model (LLM) agents are increasingly used in many applications, raising concerns about their safety. While previous work has shown that LLMs can deceive in controlled tasks, less is known about their ability to deceive using natural language in social contexts. In this paper, we study deception in the Social Deduction Game (SDG) Mafia, where success is dependent on deceiving others through conversation. Unlike previous SDG studies, we use an asynchronous multi-agent framework which better simulates realistic social contexts. We simulate 35 Mafia games with GPT-4o LLM agents. We then create a Mafia Detector using GPT-4-Turbo to analyze game transcripts without player role information to predict the mafia players. We use prediction accuracy as a surrogate marker for deception quality. We compare this prediction accuracy to that of 28 human games and a random baseline. Results show that the Mafia Detector's mafia prediction accuracy is lower on LLM games than on human games. The result is consistent regardless of the game days and the number of mafias detected. This indicates that LLMs blend in better and thus deceive more effectively. We also release a dataset of LLM Mafia transcripts to support future research. Our findings underscore both the sophistication and risks of LLM deception in social contexts.

Read the original paper