Skip to content
AI.info

Research

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

Overview Research area: Interactive language agents, user simulation, and multi-turn reinforcement learning (cs.AI / cs.CL). Technical level: Advanced. The paper assumes familiarity with supervised fi

arXiv
2610.09484
Published
2026-10-07
Authors
Hoang Phan, Dat Huynh, Andrey Zhmoginov, Qi Zeng, Wancen Mu, Yue Cao, Shengjie Bi, Yun He, Changdae Oh, Deren Lei

AI summary

Overview

Research area: Interactive language agents, user simulation, and multi-turn reinforcement learning (cs.AI / cs.CL).

Technical level: Advanced. The paper assumes familiarity with supervised fine-tuning, GRPO-style policy optimization, on-policy distillation, and multi-turn agent benchmarks.

Scope: This paper introduces MIMESIS, a purpose-built 4B/9B user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns, and shows that freezing it as a training environment improves downstream agent generalization compared with using GPT-5.5 as the simulated user.

What This Paper Is About

Training and evaluating interactive language agents requires realistic user behavior, but collecting human interactions is expensive and hard to scale. The obvious substitute — prompting helpful assistant LLMs to role-play users — distorts the task, because those models are optimized to be cooperative and explicit rather than ambiguous, impatient, or withholding. The paper's goal is to build a user simulator that is both behaviorally faithful to real users and useful as a frozen environment for multi-turn agent reinforcement learning.

Key Contributions

  1. Diagnosis of a behavioral mismatch. The authors show that off-the-shelf LLM users distort measured agent difficulty, and propose a two-stage pipeline that first learns a realistic simulator and then uses the frozen simulator to post-train an agent.
  2. MIMESIS, a 4B and 9B purpose-built user simulator. It combines user-side mid-training on human conversations, ThoughtTrace reasoning supervision, and joint multi-domain RL with a realistic-behavior objective built from 13 recurring interaction patterns derived from real user conversations.
  3. A thorough simulator evaluation. MIMESIS-9B surpasses frontier models on SOUL-Index, RealUserSim, tau-USI, and SimulatorArena, with ablations establishing the contribution of reasoning supervision and realistic-behavior training.
  4. Coached On-Policy Self-Distillation (CSD). A method that converts a simulator's private thoughts and subsequent utterances into retrospective coaching notes, which a feedback-conditioned teacher turns into dense token-level supervision for the agent — while the deployed agent conditions only on public dialogue.

Main Findings

  • Frontier users make tasks too easy; pretrained models make them too hard. On tau-bench, a fixed GPT-5.5 agent succeeds on 63.6% of interactions with real users, but three frontier API simulators raise that success rate to 82.4–84.4%, while pretrained Qwen3.5-4B and Qwen3.5-9B reduce it by 16.0 and 14.9 percentage points respectively. Every trained simulator deviates less than every off-the-shelf model, with at most 10.3 points versus at least 14.9.
  • MIMESIS-9B leads on SOUL-Index. It reaches an overall score of 65.7 versus Claude-Opus-5 (64.9), GPT-5.5 (64.1), and the strongest released simulator Osim-8B (58.8). MIMESIS-4B scores 63.7. Relative to their pretrained backbones, the 9B and 4B models improve by 15.2 and 18.0 points.
  • The largest gain is conversational simulation. MIMESIS-9B and MIMESIS-4B score 75.1 and 74.4 on the CONV axis, exceeding the strongest baseline on that axis, Ditto-8B (61.7), by 13.4 and 12.7 points. MIMESIS-9B also has the highest reported means on COG (86.2) and EVAL (69.7); GPT-5.5 remains highest on SS (70.9) and Claude-Opus-5 on ROLE (64.6).
  • Trajectory-level fidelity improves substantially. On RealUserSim PT3, MIMESIS-9B reaches a Fidelity Index of 94.0, exceeding the strongest baseline, Claude-Opus-5 (80.6), by 13.4 points; MIMESIS-4B scores 89.7. Gains concentrate in interaction and information flow (91.3 versus 69.5) and pacing (91.2 versus 72.5), while persona agreement is near ceiling for several models, including GPT-5.5 (99.3) and MIMESIS-9B (99.4).
  • Lower Turing distance on SimulatorArena. Compared with Claude-Opus-5, the strongest baseline there, MIMESIS reduces Turing distance by 3.6 points.
  • Training against MIMESIS generalizes better than training against GPT-5.5. Across eight Gym environments (three held out from training) and nine unseen evaluation user models, replacing GPT-5.5 with MIMESIS-9B under fixed GRPO raises the mean score from 26.10 to 29.54, improving performance under every evaluation user. Gains range from 1.46 points under Claude-Opus-5 to 4.55 points under Gemini-3.8-Flash.
  • CSD adds further gains on top. With the same frozen MIMESIS-9B simulator, CSD raises the mean score from 29.54 to 31.09 across all nine evaluation users. Over 100 optimization steps, the CSD curve stays above GRPO on validation reward after the earliest checkpoints, isolating the effect of the coaching objective from simulator choice.
  • Reasoning supervision matters. Mid-training on user utterances alone provides no target for the reasoning preceding them and may suppress reasoning-trace generation; ThoughtTrace supplies users' self-reported motivations as supervision for a private reasoning trace before the public utterance.

Methodology in Plain English

The approach has two decoupled stages.

Stage I — learn the simulator. The authors start from Qwen3.5 backbones at 4B and 9B and adapt them to the user role by mid-training on human–assistant conversations, predicting the human's next utterance from preceding dialogue. The training mixture contains 21.2M examples from 62 corpora, and mid-training processes 9.51B tokens. Because ordinary logs record what users say but not why, a second supervised stage uses ThoughtTrace annotations: for annotated examples, the model is supervised to produce a private reasoning trace before the public utterance, using users' self-reported motivations and interpretations of assistant responses. Hyperparameters reported include a mid-training learning rate of 3×10⁻⁵ and a ThoughtTrace SFT learning rate of 1×10⁻⁷ with cosine annealing.

The model is then optimized with GRPO across the SOUL environments used by OdysSim, but — unlike OdysSim, which trains a separate RL expert per task and distills selected trajectories — the shared simulator is trained directly on all domains, with every domain's rollouts updating the same parameters. The mixture adds a realistic-behavior domain built from 13 behavior categories derived from ThoughtTrace, including hidden evaluation criteria, incremental goalpost shifting, and clarification noncooperation. These behaviors are instantiated in ABCD customer-support scenarios with structured customer records and policies; each rollout is scored for expression of the target behavior, temporal placement, naturalness, and task consistency, with a penalty for responses that describe the behavior instead of expressing it.

Stage II — learn the agent. The simulator is frozen and used as the interactive environment. The agent is optimized with GRPO over multi-turn conversations following the UserRL extension, where task rewards produce group-normalized advantages. On top of this, CSD takes each sampled agent response, the simulator's private thought, and the next user utterance, and has a coach model produce a short improvement note. The same sampled response is then scored twice: by a student conditioned only on public history, and by a stop-gradient teacher conditioned additionally on the coaching note. The difference in log-probabilities, passed through a sigmoid with sensitivity parameter β, becomes a per-token weight in a weighted likelihood objective added to the GRPO loss with coefficient α. Tokens with negative log-probability difference get weaker reinforcement rather than a penalty, and no replacement response is decoded. Privileged information is training-only: at deployment the agent sees only the observable conversation.

Why This Matters

Impact on research. The paper reframes how user simulators should be judged — not just by surface naturalness, but by whether they preserve the interaction difficulty induced by real users and whether policies trained against them transfer to other user models. It also introduces CSD, a distinct variant of feedback-conditioned self-distillation that re-weights already sampled responses rather than decoding refinements.

Real-world applications:

  • Customer-support agent training. Simulated users that withhold information, shift goals, or refuse to clarify produce agents better prepared for messy real tickets, as demonstrated in the ABCD support scenarios.
  • Agent evaluation and regression testing. A calibrated simulator gives a reproducible, scalable stand-in for human raters when measuring multi-turn competence.
  • Tool-using assistants and troubleshooting bots. Agents that must resolve ambiguity across extended stateful interactions benefit from training against less cooperative users.
  • Pre-deployment robustness testing. Simulators spanning many behavior patterns can stress-test policies before they reach production users.

Industry relevance. Large-scale agent post-training needs environments that scale, and human data does not. Demonstrating that a 9B open-weight simulator outperforms frontier API models on simulation benchmarks — and that training against it transfers better than training against GPT-5.5 — suggests a cheaper, controllable, and more realistic alternative for agent development pipelines. The paper reports a public code release.

Future Directions

  • Human-validated transfer. The agent evaluation uses nine simulated evaluation users; the paper does not report a study with human participants checking whether the transfer gains hold with real users.
  • Scaling the simulator. Only 4B and 9B variants are reported, leaving open how simulator quality and downstream agent gains scale with size and data.
  • Extending the behavior taxonomy. The 13 patterns are derived from recurring patterns in ThoughtTrace; whether they cover the full space of user difficulty, and how they generalize to new domains beyond ABCD scenarios, is not established.
  • Coaching design and cost. The coach model, the sensitivity parameter β, and the loss weight α are not explored in the reported results, leaving the sensitivity of CSD to its design choices as an open question.

Target Audience

Researchers and engineers working on interactive language agents, multi-turn reinforcement learning, and dialogue simulation. It will be most useful to readers already comfortable with policy-gradient methods and self-distillation, and to practitioners who need realistic, scalable user environments for post-training or evaluating agents rather than single-turn benchmark scores.

Authors’ abstract

Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, most agent-training frameworks rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real users. We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions. Empirically, our 9B model achieves a SOUL-Index of 65.7, surpassing the strongest frontier model. Compared with Claude-Opus-5, the strongest baseline on RealUserSim and SimulatorArena, MIMESIS improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points, respectively. We then freeze the simulator and train an agent by interacting with the frozen simulator using multi-turn reinforcement learning. Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators, demonstrating stronger generalization to new user simulators. Moreover, we propose Coached On-Policy Self-Distillation (CSD), which leverages simulator-generated private reasoning traces and subsequent utterances as feedback on how well the agent addresses user needs. A coach converts this information into concise coaching notes that describe how the agent can better anticipate user needs and adapt its behavior over the course of an interaction. CSD turns this feedback into dense, token-level supervision beyond sparse task rewards, yielding further gains across all nine evaluation user models.

Read the original paper