Skip to content
AI.info

Research

Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

Overview Research area: AI alignment and safety, specifically strategic deception by LLM/VLM agents in multi-agent embodied environments (social-deduction games, multimodal Minecraft sandboxes, agent

Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
arXiv
2608.30428
Published
2026-08-31
Authors
Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira, Gunhee Kim

AI summary

Overview

Research area: AI alignment and safety, specifically strategic deception by LLM/VLM agents in multi-agent embodied environments (social-deduction games, multimodal Minecraft sandboxes, agent harness design).

Technical level: Advanced. The paper assumes familiarity with LLM/VLM agent architectures, ablation methodology, and correlation statistics, and it defines a new game environment, an agent harness, and a 23-label annotation taxonomy.

Scope: The paper builds a 3D multimodal Among Us sandbox plus a configurable VLM agent harness, and uses 192 harness-ablation matches and 1,152 cross-model matches to test whether imposter agents deceive through verbal channels, non-verbal channels, or both.

What This Paper Is About

Existing testbeds for studying deceptive AI agents are text-only social-deduction games (Mafia, Werewolf, Avalon, Among Us) run under a single fixed agent configuration, so they cannot exercise non-verbal channels such as physical action, visual perception, and spatial behavior, and they cannot tell whether an observed behavior comes from the underlying model or from the surrounding agent scaffold. The authors build a 3D embodied Among Us sandbox in Minecraft where imposter agents must deceive crewmates through joint verbal and non-verbal action, plus a configurable agent harness whose components can be independently ablated. The goal is to attribute deception behaviors to specific cognitive components or specific VLM backbones rather than to one chosen setup.

Key Contributions

  1. MineAmongUs, described as the first 3D multimodal Among Us sandbox in Minecraft, where VLM agents perceive through raw egocentric RGB observations (360 × 640), navigate a continuous physical space, physically interact with objects and other agents, and chat in natural language during meetings. This jointly exposes perceptual, spatial, and conversational deception channels in a single match.

  2. Aria (Ablation-Ready Imposter/crewmate Agent), a configurable VLM agent harness exposing five cognitive components as independent ablation axes: state representation, memory, planning, reflection with skill memory, and prompt style. The authors frame it as a controlled instrument for attribution rather than an autonomy-maximizing agent system.

  3. A 23-atom annotation scheme for deception, derived from independent per-step annotation of 48 gameplay logs by two authors and grounded in three deception taxonomies (Whiten–Byrne, Whaley, and Buller & Burgoon's Interpersonal Deception Theory). The atoms are organized into a two-level hierarchy of six clusters (three non-verbal, three verbal), plus a higher-order "arc" defined as a multi-step sequence of atoms realizing one deceptive intent.

  4. A validated LLM-as-a-Judge pipeline based on Qwen3.6-27B that reaches near-human atom-labeling agreement (human–human Cohen κ = 0.792 and Spearman ρ = 0.803; human–LLM Cohen κ = 0.709, F1_w = 0.792, Spearman ρ = 0.713), enabling annotation at scale.

Main Findings

  • Egocentric state collapses gameplay. With imposters restricted to ego (first-person) state across 10 games (5 per VLM backbone), imposters landed 0 kills, and imposter win rate dropped from 40% to 0% for GPT-4.1-mini and from 40% to 20% for Qwen3.6-27B (the single Qwen-ego win was a default timeout with 0 kills). A complementary pilot restricting crewmates to ego confirmed the symmetric collapse, so all remaining experiments fixed state to privileged for both roles.

  • Harness composition shifts imposter win rate even with fixed VLM backbones. Varying only the crewmate-side (memory, planning) slice raised imposter WR by +8 pp under Qwen3.6-27B and dropped it by −35 pp under GPT-4.1-mini. Cell-level WR across 48 matches per cell: 1-1 (Qwen × Qwen, semantic/reactive) 44%, 1-2 (Qwen × Qwen, window/hierarchical) 52%, 2-1 (mini × mini, semantic/reactive) 60%, 2-2 (mini × mini, window/hierarchical) 25%.

  • Cognitive axes show directional but mostly non-significant trends. A pooled two-proportion z-test over all 192 games found planning and reflection-skill reached marginal significance in the pre-registered direction (Δ ≈ +9.4 pp, p ≈ 0.096 at α = 0.10); memory and prompt were not significant. Directionally, window memory tended to beat explicit LLM-based belief updating, hierarchical planning tended to beat reactive planning, and post-meeting reflection was generally beneficial, while prompt style varied more strongly by VLM backbone.

  • Non-verbal kill-cycle atoms carry the strongest win associations in the ablation. Point-biserial correlations over all 192 games were strongest for Witness-Aware Kill (r_pb = +0.434, 618 counts), Post-Kill Flee (r_pb = +0.414, 352 counts), Strategic Non-Reporting (r_pb = +0.270, 187 counts), and Stalking Pre-Kill (r_pb = +0.208, 1,803 counts). These cover the stalk → witness-aware kill → flee loop and deliberate non-reporting of bodies.

  • Different axes win through different channels. For memory, reflection-skill, and prompt, the higher-WR settings (window, meeting-on, minimal) produced more of the non-verbal atoms E, F, H, C. The planning axis behaved differently: hierarchical produced substantially more verbal V-1 Falsification atoms (P, O, M, S, R, N), whose correlations with imposter wins were positive but more moderate (r_pb 0.137 to 0.210).

  • VLM backbone predicts performance across both roles. Across the 12 evaluated VLM backbones, imposter WR and crewmate WR correlated within a model at Pearson r = +0.71, indicating the underlying backbone matters more than role-specific capability.

  • Winners are distinguished most by non-verbal camouflage. Among the 6 selected VLMs (576 judged matches), Fake-Mission Performance showed the largest winner–loser gap: top-3 models produced 6.6× more per game than worst-3 (12.64 vs. 1.91, r = +0.72). Other winner atoms were Post-Kill Flee (r = +0.59, 1.2×) and Witness-Aware Kill (r = +0.52, 2.7×), plus verbal atom Pattern Fabrication (r = +0.49, 1.5×). Loser atoms split into defensive verbal reactions (Counter-Accusation r = −0.52; Honesty/Credibility Marker r = −0.50; Throw-Under-Bus r = −0.54) and passive non-verbal posturing (Blend-In Wandering r = −0.62; Bystander Co-flight r = −0.58).

  • Cluster-level correlations split into positive and negative groups. At the cluster level, NV-1 Camouflage was the strongest positive signal (r = +0.54), followed by NV-3 Report & Emergency (r = +0.44), while NV-2 Pursuit & Kill (r = −0.37), V-1 Falsification (r = −0.42), V-2 Equivocation (r = −0.52), and V-3 Concealment (r = −0.48) were negative. The authors note the small cross-model sample (N = 6) means sign and relative magnitude should be used for comparison rather than statistical inference.

  • No single winning strategy. The top-3 VLMs reached high imposter WR through two distinct archetypes: a non-verbal Camouflage-heavy route (Gemini-3-flash with NV-1 share 39.7% and 70.8% WR; Gemini-3.1-flash-lite with NV-1 share 38.4% and 65.6% WR) and a verbal Falsification-heavy route (Kimi-K2.5 with V-1 share 43.2% and 66.7% WR).

Methodology in Plain English

The researchers picked Among Us because it is a hidden-role game that already separates two phases: a task phase where players move and act physically, and a meeting phase where they talk and vote. That structure lets deception happen in two different channels within one match. They rebuilt the game in Minecraft with 8 players (2 imposters, 6 crewmates), giving agents an egocentric RGB camera view, scoreboard signals, chat history, and their own position, and letting them act by dispatching Mineflayer JavaScript programs (sequences of primitives such as pathfinding, attacking, and fleeing). Matches run until one of four win conditions is met: crewmates alive drops to or below imposters alive, the match exceeds the default 200-step budget, all imposters are voted out, or all missions are completed (3 per crewmate, with ghosts of dead crewmates still contributing). Meetings are turn-based with each alive player speaking up to 3 turns.

Their agent harness, Aria, runs a planner that dispatches one of eight decision modules per step (kill, report, surveillance, emergency, meeting, vote, move, mission; the last two of the eight, surveillance and emergency, are auto-triggered for crewmates only). Each module combines a VLM decision with generation of an action program. Imposter actions carry a private one-line narration of intent that other agents never see, which is later used as evidence for annotation.

To make findings attributable, they turned five components into switchable settings: state representation (ego versus privileged full-map view), memory (flat rolling window versus semantic per-player belief tracking by a dedicated LLM), planning (reactive single-step versus hierarchical long-horizon), reflection and skill memory (off versus meeting-end reflection with a skill buffer), and prompt style (deterministic tactical prompt versus minimal rules-only prompt).

For measurement they had two authors annotate 48 gameplay logs, yielding 23 atoms and arcs, and then verified that an LLM judge reproduced human labels closely enough to annotate at scale. Experiment 1 varied the harness (16 imposter configurations × 2 crewmate styles × 2 VLM pairings × 3 repetitions = 192 matches) while holding the VLM fixed. Experiment 2 ran a 12 × 12 round-robin over 12 VLM backbones under 4 fixed harness combinations with 2 repetitions each (1,152 matches).

Why This Matters

This work moves AI deception research from text-only, single-configuration testbeds to an embodied multimodal setting where agents can lie with their bodies as well as their words — and where a lie like pretending to do a mission is directly visible to any observer. Crucially, it shows that the surrounding agent harness, not just the model, can move win rates by tens of percentage points, which means many prior findings about "model deception" may be entangled with configuration choices. It also supplies an annotation vocabulary and a validated automated judge that other researchers can reuse.

Real-world applications:

  • Safety evaluation for deployed agents: The harness-ablation logic offers a template for testing whether an agent's behavior is a property of the model or of its scaffolding before deployment decisions are made.
  • Multi-agent coordination and communication: The atom taxonomy covers tactics like equivocation, concealment, and counter-accusation that appear in negotiation, bidding, and customer-service agents operating alongside other agents.
  • Robotics and human-robot interaction: The finding that non-verbal channels dominate in embodied play is directly relevant to robots whose physical behavior (positioning, timing, task performance) is observable evidence to humans.
  • Game AI and interactive entertainment: The sandbox and annotation scheme can be reused to design NPCs whose deception is legible and tunable, or to build detection systems for deceptive agents in online multiplayer games.

Industry relevance is highest for teams building autonomous agents that act on a user's behalf and interact with other agents, and for anyone whose evaluation pipeline currently relies on a single fixed agent configuration or on win-rate alone.

Future Directions

  • Testing whether atoms transfer to open-ended interaction. The authors explicitly state that their measurements should be read as game-bounded behavioral deception, and that establishing how these behaviors transfer to open-ended human interaction remains future work. They note that atoms like Strategic Non-Reporting, Self-Reporting Kill, Weaponized Meeting, Alibi Fabrication, and Fake Eyewitness Testimony target other agents' beliefs more directly than task execution does.

  • Closing the egocentric perception gap. Because both roles had to be fixed to privileged state for viable play, and because neither local text tracking alone nor privileged text without RGB sustained viable play, a clear next step is improving VLM spatial localization so realistic entangled, first-person perception becomes testable.

  • Scaling the cross-model comparison. The winner–loser atom analysis used only 6 of the 12 VLMs (N = 6), which the authors flag as too small for statistical inference; a broader model pool would let the reported correlations be tested rather than compared.

  • Understanding the two winning archetypes. Since top models won either through camouflage-heavy or falsification-heavy routes, an open question is what determines which strategy a given backbone adopts, and whether mixing opponents of different archetypes changes outcomes.

  • Arc-level analysis. The paper defers the arc-level view — sequences of atoms forming a multi-step deceptive intent — to an appendix, leaving multi-step deception dynamics as a direction for further investigation.

Target Audience

AI alignment and safety researchers studying deceptive or strategic behavior in LLM/VLM agents; multi-agent systems researchers who need configurable, attributable testbeds; embodied AI and robotics researchers interested in nonverbal behavioral signals; and game AI or interactive-entertainment developers building deceptive NPCs or deception-detection systems. Practitioners running agent evaluations will also benefit from the harness-ablation framing, though the paper is written for a research audience comfortable with ablation grids, correlation statistics, and VLM agent architectures.

Authors’ abstract

Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction games (where each player holds a hidden role and communicates with others to deduce identities) serve as the canonical testbed, particularly in multi-agent settings. Existing testbeds, however, are text-only and run on a single fixed agent configuration, missing the non-verbal sensorimotor channels treated as core by deception taxonomies and leaving it ambiguous whether an observed behavior reflects the underlying model or the surrounding harness. We introduce MineAmongUs, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action. We also propose ARIA, a configurable VLM-agent harness that exposes five cognitive-component ablation axes; and an atom- and arc-level annotation scheme grounded in deception taxonomies and operationalized at scale by an LLM-as-a-Judge reaching near-human atom-labeling agreement. Empirical results show that VLM agents pursue imposter wins through joint verbal and non-verbal deception, with non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation. Taken together, our work opens a new path for embodied VLM-agent alignment research.

Read the original paper