Skip to content
AI.info

Research

HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models

Overview Research area: Full-duplex speech language models (SLMs) and reinforcement learning for real-time spoken dialogue, specifically credit assignment for turn-taking, backchanneling, and interrup

HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models
arXiv
2610.07727
Published
2026-10-06
Authors
Kyudan Jung, Hyunsin Park, Yoonhyung Lee, Jinhwan Park, Jinhyeok Yang, KiHyun Nam, Jaegul Choo, Jinkyu Lee

AI summary

Overview

Research area: Full-duplex speech language models (SLMs) and reinforcement learning for real-time spoken dialogue, specifically credit assignment for turn-taking, backchanneling, and interruption handling.

Technical level: Advanced. The paper assumes familiarity with policy-gradient RL (GRPO/PPO), streaming speech tokenization, and full-duplex architectures such as Moshi and PersonaPlex.

Scope in one sentence: HiPLEX factorizes a pretrained full-duplex text policy into a "when to speak" control policy and a "what to say" content policy, routing timing rewards through event-causal masks and semantic rewards through an LLM judge.

What This Paper Is About

Full-duplex speech models must decide not just what to say but when to say it, how long to speak, and when to yield the floor. Existing RL approaches for this setting either push timing feedback onto every token in a fixed window around an annotated event, or optimize semantics while leaving timing alone, so the two problems are never improved jointly with correctly placed credit. The paper's goal is a post-training method that assigns timing feedback to the decisions that actually caused a timing outcome, and semantic feedback to the content decisions, without adding new model parameters.

Key Contributions

  1. A two-level conditional action hierarchy with no architectural change. The pretrained text policy is factored into a high-level decision over whether and when to invoke content emission (choosing among pad, epad, and cont) and, only when cont is selected, a low-level decision over which content token to emit. Both factors come from the same pretrained head and introduce no new parameters.

  2. Credit that follows event structure rather than proximity. Event-causal credit masks are sign-dependent and derived from the policy's own generated speech episodes: a late response penalizes the waiting decisions that delayed it, while a timely response credits its onset. A fixed window applies one advantage to every token inside it and cannot make this distinction.

  3. A post-training method for frontier full-duplex models with a multi-axis analysis. On Full-Duplex-Bench v1, HiPLEX gives the strongest overall trade-off on Moshi and improves the restraint axes on PersonaPlex. The authors report a post-boundary word-activity diagnostic for the mixed PersonaPlex turn results and analyze which interaction styles different metrics favor.

  4. Diagnosis of training sensitivity across model families. The same recipe collapses on PersonaPlex at Moshi's learning rate, requiring a rebalanced reward configuration, which the authors attribute to PersonaPlex's more talkative, floor-holding prompt prior.

Main Findings

  • Pause and backchannel restraint improve on Moshi. On Full-Duplex-Bench v1, HiPLEX reaches a CANDOR pause takeover rate of 0.306 versus 0.454 for the reproduced GRPO baseline, and a backchannel takeover rate of 0.127 versus 0.309. These gains hold across three Moshi seeds.

  • Post-interruption response latency shortens on Moshi. HiPLEX records 0.439 on this axis versus 0.660 for GRPO, while judged interruption-response quality remains comparable (4.083 versus 3.942 for the displayed seed-42 checkpoint).

  • PersonaPlex gains on restraint come with mixed turn results. HiPLEX reduces pause takeover (0.218 versus 0.444 for GRPO) and backchannel takeover (0.018 versus 0.236), but also shows a lower official smooth-turn TOR (0.714 versus 0.882) and longer post-interruption latency (0.590 versus 0.548). A complementary diagnostic shows more qualifying post-boundary word activity and a shorter median offset to the first post-boundary word in both model families, and PersonaPlex begins speaking before the boundary less often.

  • The same recipe needs a different step size on the two families. At Moshi's 4×10⁻⁶ learning rate, PersonaPlex training collapses with suppression rising while initiation declines; halving the rate only delays convergence to the same near-silent behavior. Rebalancing rewards toward initiation prevents the collapse. Reported settings are 4×10⁻⁶ for HiPLEX and 1×10⁻⁶ for the flat GRPO baseline on Moshi, and 2×10⁻⁶ for PersonaPlex.

  • HiPLEX better matches human timing marginals than GRPO. Wasserstein-1 distance to pooled Seamless–Fisher marginals on Moshi: turn taking 0.902 versus 1.781 for GRPO, backchannel rate 0.327 versus 1.512, and backchannel length 0.238 versus 0.266. GRPO is closer only on pause intrusion (0.012 versus 0.080), because it is nearly always silent during hesitations while the human rate is small but nonzero. PersonaPlex remains mixed: HiPLEX is closer on turn timing (2.627 versus 3.511) and backchannel rate (1.187 versus 1.652) but not on the other two marginals.

  • Semantic feedback helps HiPLEX more than GRPO. In a matched 100-epoch ablation, adding the LLM-judge reward substantially improves judged interruption-response quality for HiPLEX (2.332 without versus 4.083 with), while GRPO shows only a small gain (3.859 to 3.942). For HiPLEX, semantic feedback also reduces pause and backchannel takeover rates and shortens post-interruption latency.

  • Event-causal masking is the credit rule that works. Holding the factorized policy and rewards fixed, all-frame credit lowers the turn-response rate (0.958) and raises post-interruption latency (0.874); a fixed-window variant reaches 0.933 turn rate and 0.465 latency; matched-random sparsity collapses toward silence, answering only 6.7% under the official turn-taking rule. Event-causal routing is the only variant reaching an official turn TOR of 1.0 while remaining competitive on pause handling (0.306), post-interruption latency (0.439), and judged quality (4.083). Random masking is excluded from ranking because its 0.067 turn rate indicates silence.

  • Masking is sparse by construction. Event-causal masking selects approximately 20–30% of the frames used by a flat update in the reported runs.

Methodology in Plain English

The model streams audio and an aligned text stream on an 80 ms frame clock (12.5 Hz for Moshi). Rather than treating every text token as an independent action, the authors split the existing text distribution into two decisions at each frame. The first decision, the control policy, uses log-sum-exp to compute the total probability mass of the pad, epad, and cont token groups and then softmaxes over the three groups. The second decision, the content policy, is the ordinary conditional distribution inside the cont group and is only evaluated when cont was chosen. Multiplying the two factors exactly reconstructs the original text-token probability, so nothing is added to the model.

Training is group-based. For each annotated interaction window, the method samples a group of rollouts, then standardizes each timing reward component independently within that group, following GRPO-style group-relative advantages. Reward components differ by interaction type: three for turn taking, two for pause handling, five for backchanneling, and three for interruption handling. The interruption axis, for instance, penalizes speech that overlaps the user's interruption after a 0.16 s grace period.

The distinctive step is deciding which frames receive which advantage. The model's own generated audio is passed through Silero VAD; consecutive speech spans separated by gaps of at most 1 s are merged into speech episodes, with episodes of at most 1 s labeled short (candidate backchannels) and longer ones sustained (turn and re-entry onsets). A binary mask is then set to one only at the control decisions causally responsible for that component, with the mask's behavior depending on the sign of the advantage. A penalty for a late response targets the preceding waiting interval and leaves the eventual onset untouched; a positive onset advantage credits the decision to begin speaking; failure-to-yield penalties target speech overlapping the user. When components select the same frame, their weighted advantages add.

Semantics are handled separately. The user context and generated speech are transcribed with Parakeet-TDT-0.6B-v2, and Gemini 2.5 Flash-Lite assigns a response-level score in {0, 1, 2}. This score is standardized within the rollout group and applied only to the content policy at frames where a content token was generated. Both policies are updated with PPO-style clipped losses, each normalized by its own count of selected decisions, with separate KL regularization toward a frozen reference model. Training uses conversation-only subsets of the Seamless-Interaction corpus, from which the authors extract 73,561 annotated interaction windows roughly balanced across the four conversational behaviors. Evaluation uses Full-Duplex-Bench v1 with the official streaming protocol, real-time playback from t = 0, and no silence modification; response quality is scored by gpt-4o-2024-08-06 on frozen ASR transcripts.

Why This Matters

Impact on research. The paper reframes full-duplex speech RL as a hierarchical credit-assignment problem rather than a flat token-level one, and shows empirically that factorization alone or sparsity alone is not sufficient: the routing rule is what produces the result. It also argues that event-level success metrics can obscure early speech and excessive silence, and that evaluation against human timing distributions is needed alongside binary benchmarks.

Real-world applications:

  • Voice assistants and conversational agents that must stop speaking when a user interrupts and resume appropriately.
  • Customer-service and call-center agents where staying silent during a customer's hesitation and acknowledging without taking the floor are core quality requirements.
  • Telephony or on-device conversational systems where adding no new parameters and reusing the existing text head keeps deployment cost unchanged.
  • Turn-aware interfaces such as accessibility tools or tutoring systems, where mistimed speech is disruptive regardless of whether the answer is correct.

Industry relevance. The authors are affiliated with Qualcomm AI Research, and the method's no-new-parameters property, its separate learning-rate tuning per model family, and its explicit treatment of training collapse under an aggressive learning rate are practical engineering concerns for teams deploying streaming speech models.

Future Directions

  • Generalizing the factorized hierarchy beyond Moshi and PersonaPlex. The paper shows that PersonaPlex's prompt prior changes the optimization margin substantially, but does not report whether the rebalanced reward settings transfer to other duplex backbones or to additional languages.
  • Reconciling the mixed PersonaPlex turn results. HiPLEX improves restraint and post-boundary word activity on PersonaPlex yet lowers the official smooth-turn TOR and lengthens post-interruption latency; the paper reports this discrepancy and a diagnostic but does not resolve which axis better reflects conversational quality.
  • Improving evaluation. The paper argues that binary event metrics obscure early speech and excessive silence and that human timing marginals should be compared, leaving open the question of which combined metric should become standard.
  • Understanding rubric and feedback strength. The authors note that rubric choice and semantic-reward strength shift the balance between judged quality and participation, with stronger settings sometimes raising conditional quality as response rates fall; how to set these trade-offs is not resolved.

Target Audience

Researchers and engineers working on full-duplex or streaming speech language models, reinforcement learning for dialogue and turn-taking, or post-training of audio-LLMs. It will also interest practitioners who need turn-taking behavior in deployed voice agents and readers following credit-assignment methods for sequential decision problems.

Authors’ abstract

As human--AI interactions become more conversational, full-duplex speech language models capable of natural real-time dialogue are growing in importance. Beyond generating appropriate responses, these models must coordinate turn-taking, backchanneling, and floor management in real time. Reinforcement learning (RL) provides a way to refine these behaviors through direct feedback on interaction outcomes. However, existing RL methods either apply timing feedback to a token policy or optimize semantic content, leaving the joint improvement of timing and content unresolved. We introduce HiPLEX, an RL framework that factorizes a pretrained full-duplex text policy into a control policy that decides when to emit content and a conditional content policy that decides what to emit. The first factor selects among 'pad', 'epad', and 'con'. The second selects a token only when 'con' is chosen. This hierarchy describes conditional actions within each frame and uses the model's existing text head. We route timing advantages to the token-group factor through event-causal masks derived from generated speech episodes, and route an LLM-judge semantic advantage to the conditional content factor. Across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities, and shortens post-interruption response latency relative to GRPO, while maintaining comparable judged interruption-response quality. On Moshi and PersonaPlex, HiPLEX better matches pooled human turn-timing and backchannel-rate marginals than GRPO.

Read the original paper