Skip to content
AI.info

Research

Learning Steerable Clarification Policies with Collaborative Self-play

Overview Research area: Conversational AI / dialogue grounding and ambiguity resolution; reinforcement learning post-training for large language models. Technical level: Intermediate. The paper assume

Learning Steerable Clarification Policies with Collaborative Self-play
arXiv
2512.04068
Published
2025-12-03
Authors
Jonathan Berant, Maximillian Chen, Adam Fisch, Reza Aghajani, Fantine Huot, Mirella Lapata, Jacob Eisenstein

AI summary

Overview

Research area: Conversational AI / dialogue grounding and ambiguity resolution; reinforcement learning post-training for large language models.

Technical level: Intermediate. The paper assumes familiarity with post-training methods (reinforcement learning from rollouts, Reinforced Self-Training) and with self-play as a training paradigm, but its central idea — deciding between guessing, enumerating, or clarifying — is explained concretely.

Scope: The paper proposes training a single "steerable" assistant that conditions on numeric cost coefficients supplied in its prompt, so that the same model will ask clarification questions, enumerate multiple interpretations, or answer directly depending on how expensive each option is.

What This Paper Is About

When a user's query is ambiguous, an AI assistant has three broad options: guess a likely intent and answer, enumerate several possible intents and answer each, or ask a clarifying question. No option is always best — it depends on context such as user preferences, screen size, or whether the user is driving. This paper trains an assistant that is steerable: it receives numerical costs for clarifications and for generating words, and learns to pick the conversational action that maximizes accuracy penalized by those costs. The training uses collaborative self-play, where one language model plays the user and another plays the assistant.

Key Contributions

  1. A steerable formulation of grounding behavior. Rather than assuming ambiguity should always be resolved by clarification, the paper defines a reward that trades off accuracy against the number of clarification turns and the length of the final answer, parameterized by scalar coefficients α (cost of a clarification question) and β (per-token cost). The coefficients are given to the model as input.

  2. A collaborative self-play training procedure with action-sequence filtering. The authors apply Reinforced Self-Training (ReST), sampling conversations between a fixed user simulator and a trainable assistant across many cost-coefficient settings. To avoid training on lucky guesses, they identify the action sequence with the highest expected reward across interpretations, then keep the best rollout compatible with that sequence for each interpretation.

  3. Empirical evidence of steerability where strong prompted models fail. On two QA benchmarks, the trained Steerable Grounding Policy (SGP) achieves the highest F1 and reward among the compared methods and adjusts its clarification and multi-answer rates in the correct direction as α and β change; prompted baselines, including chain-of-thought prompting, do not.

  4. Generalization to unseen cost coefficients. SGP smoothly interpolates to α and β values never used during training, extending user control beyond the training grid.

Main Findings

  • SGP achieves the best reward and F1 among the compared methods. On AmbigQA, SGP reaches a reward of 6.63 and F1 of 19.94 on ambiguous queries, and 31.36 reward / 41.28 F1 on unambiguous queries, versus 1.32 / 14.33 and 26.73 / 35.58 for the constrained Answer baseline and −34.49 / 16.16 and −15.01 / 33.07 for the Prompted baseline.

  • The same holds on Pacific. SGP reaches 54.32 reward / 65.73 F1 on ambiguous queries and 65.70 / 79.52 on unambiguous queries, compared with 42.44 / 52.25 and 56.91 / 71.52 for Answer, and 33.00 / 50.04 and 52.07 / 68.43 for Prompted-COT.

  • Prompted models do not condition on cost coefficients. On AmbigQA, the Prompted baseline behaves counterintuitively — increasing its clarification rate as α rises and increasing its multi-answer rate and turn length as β rises. Prompted-COT does not vary its clarification or multi-answer rates with α or β at all. The authors conclude the base model cannot reason about numerical costs, whereas fine-tuning teaches this capability.

  • SGP exhibits the intended sensitivity. Higher α reduces the fraction of clarifications and higher β reduces the fraction of multi-answers, on both AmbigQA and Pacific. SGP also uses clarifications more often for ambiguous queries than unambiguous ones (for example, on Pacific the ambiguous clarification rate rises to 42.86%).

  • Steerability generalizes to unobserved coefficients. When varying α with β fixed at 5, and varying β with α fixed at 20, the fraction of clarification questions and multi-answers is monotonic in the coefficients, including at coefficient values (marked as purple stars) that never appeared in training.

  • Prompt-level steerability is partial. Measured as whether the model correctly predicts when the optimal action sequence changes, SGP outperforms baselines across the board on both datasets, but results are substantially higher for steering α than β; for β, precision is high but recall is low. The authors attribute this to the base model producing better training rollouts for clarifications than for multi-answers, and note substantial headroom remains.

  • Gemini models are not reliably steerable out of the box. On AmbigQA, Gemini-2.5-Flash's clarification rate increases with α (23.58% at α = 0.0 rising to 49.59% at α = 20.0), and it shows no sensitivity to β. Gemini-3-Flash shows moderate steerability (clarification rate falls from 33.86% to 22.42% as α rises), but overuses multi-answers (67.43%, 64.91%, 58.10% across β values), which reduces its reward. Both Gemini models have higher F1 than SGP due to stronger parametric knowledge, but lower reward on ambiguous questions, and Gemini-3 also has lower reward on unambiguous questions.

  • Fine-tuning teaches more than action selection. Because SGP is trained on both actions and observations, it also learns which clarification questions and which multi-answers are likely to yield high reward.

  • Improvements are consistent across settings. Reward improves for all nine coefficient pairs, not only on average; action distributions show predictable patterns; and reward and F1 improve monotonically across ReST epochs. The paper reports Table 2 without confidence intervals for readability but provides them in Appendix F, Table 9.

Methodology in Plain English

The setup is a conversation between two language models. A user simulator sees an ambiguous query plus one unambiguously specified interpretation, and the assistant sees only the query. The assistant can take three actions: #CLARIFY (ask a question), #ANSWER (answer directly), or #MULTI_ANS (give answers for several plausible interpretations). The user responds to clarification questions, and once it receives an answer it produces a final answer, which is scored against the gold answer using token-level F1 on a 0–100 scale.

The reward subtracts penalties from that accuracy: one penalty for each clarification question (weighted by α) and one for each word in the assistant's final answer (weighted by β). Because α and β are handed to the assistant as inputs, different coefficient values make different action sequences optimal — this is exactly what makes the policy steerable.

Training proceeds by self-play. For each query and each sampled coefficient pair, the system generates many rollouts (192 at training time, 64 at test time), varying α over {0, 2, 20} and β over {0.1, 0.7, 5.0}, and allowing at most one clarification question (yielding four possible action sequences: answer or multi-answer, with or without clarification). Invalid rollouts are filtered out: those with F1 below 0.1 on AmbigQA or below 0.4 on Pacific, and those where the user ignored the assistant's answer (detected when more than half the tokens in the user's final answer do not appear in the assistant's answer).

Rather than simply training on the single highest-reward rollout — which could reward a lucky guess — the method coarsens rollouts into action sequences, computes the expected reward of each action sequence across interpretations, selects the best sequence, and then adds the best rollout compatible with that sequence for each interpretation. Only the assistant's turns are trained on. This Reinforced Self-Training loop runs for T = 3 epochs, with the model resampled each epoch and the checkpoint chosen by lowest held-out log-likelihood on a development split of the generated data. Only the learning rate is tuned.

The base model is Gemma 2 9B, used for both the user simulator and the assistant. Comparisons include a direct prompt with examples, a chain-of-thought prompt, and several constrained variants (always answer, always multi-answer, always clarify, clarify then multi-answer), plus a fine-tuned direct-answer model and an oracle that retrospectively selects the action sequence with the highest expected reward to approximate an upper bound.

The datasets are AmbigQA, filtered model-assisted to remove noisy annotations and subsampled to a 2:1 ambiguous-to-unambiguous ratio (1,776 training and 382 development examples, averaging 2.41 interpretations per query in training and 2.76 in development), and Pacific, a multi-turn QA task grounded in financial documents and tables (3,744 training and 640 development examples, with equal numbers of ambiguous and unambiguous queries). Pacific is asymmetric: the assistant has private context the user cannot see, which prevents the user from answering from its own knowledge.

Why This Matters

Impact on research. The paper reframes ambiguity handling as a control problem rather than a fixed policy. Prior self-play work largely assumed clarification is the right response to ambiguity; this work shows that the correct grounding strategy is a function of explicit costs, and that a single model can be trained to span that spectrum. It also provides a concrete demonstration that strong prompted models — including Gemini-2.5-Flash and Gemini-3-Flash — are not reliably steerable by numerical cost signals, which is a useful negative result for anyone assuming prompting suffices for cost-aware behavior. The action-sequence filtering step is a methodological contribution to self-play training, addressing the problem of reinforcing lucky trajectories.

Real-world applications.

  • **Voice

Authors’ abstract

To handle underspecified or ambiguous queries, AI assistants need a policy for managing their uncertainty to determine (a) when to guess the user intent and answer directly, (b) when to enumerate and answer multiple possible intents, and (c) when to ask a clarifying question. However, such policies are contextually dependent on factors such as user preferences or modality. For example, enumerating multiple possible user intentions is cumbersome on small screens or in a voice setting. In this work, we propose to train steerable policies for managing this uncertainty using self-play. Given two agents, one simulating a user and the other an AI assistant, we generate conversations where the user issues a potentially ambiguous query, and the assistant needs to determine how to respond. Importantly, the model takes as input the numerical cost of each clarification question, and each generated word, and is asked to take the action that will maximize its final reward, which is the cost-penalized accuracy. We use Reinforced Self-Training (ReST) to train our model to achieve high reward and show this leads to a steerable policy that changes its behavior predictably conditioned on the provided costs, leading to higher reward and accuracy. Moreover, our procedure also generalizes to numerical cost values that were unobserved at training time.

Read the original paper