Research
EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation
Overview Research area: Emotional text-to-speech (TTS) synthesis, specifically training-free activation ("vector") steering of frozen speech language models. Technical level: Intermediate. Readers nee

- arXiv
- 2609.38157
- Published
- 2026-09-29
- Authors
- Kuan-Po Huang, Haohe Liu, Puyuan Peng, Haibin Wu, Zhaoheng Ni, Hung-yi Lee, Jinwon Lee, Neha Chachra
AI summary
Overview
Research area: Emotional text-to-speech (TTS) synthesis, specifically training-free activation ("vector") steering of frozen speech language models.
Technical level: Intermediate. Readers need some familiarity with TTS backbones, language-model internals, and the idea of adding directions to hidden activations at inference time.
Scope: The paper proposes EmoRES, a training-free reweighting of the shared and residual components of emotion steering vectors, and evaluates it on two frozen TTS backbones against the CoCoEmo steering baseline using objective metrics, component ablations, embedding visualizations, and human judgments.
What This Paper Is About
Emotion-conditioned TTS systems do not always produce speech that strongly matches the emotion a user requests, and fixing this by additional training costs both computation and emotion-labeled speech data. The paper studies vector steering, which leaves the TTS model frozen and instead modifies its internal activations at inference time. Its central observation is that the existing method CoCoEmo treats each emotion vector as one indivisible direction scaled by a single global strength, while in fact each vector splits into a shared "away from neutral" component and an emotion-specific residual. EmoRES controls these two components separately.
Key Contributions
- The authors conduct what they describe as the first functional analysis of the shared component and category-specific residuals inside emotion steering vectors for LM-based TTS.
- They introduce EmoRES (Emotion Residual-Enhanced Steering for TTS), a training-free generalization of conventional steering that reweights the shared and residual components without learning new subspaces or updating TTS model parameters.
- They show that CoCoEmo is an exact special case of their formulation: at λ_c = λ_r = 1 the centroid cancels and the residual-enhanced vector reduces to the original emotion vector v_e.
- They support the shared-residual interpretation with experiments across two TTS backbones (IndexTTS-2 and CosyVoice2), held-out out-of-distribution evaluation, component ablations, embedding probes, and human judgments.
Main Findings
- Rank correlation gains dominate: On IEMOCAP, EmoRES improves Spearman rank correlation by 26.13 percentage points for IndexTTS-2 (22.00% to 48.13%) and 12.97 points for CosyVoice2 (39.13% to 52.10%), which the paper reports as relative gains of 118.8% and 33.1%.
- Hit rate improves: Emotion hit rate rises by 12.95 points for IndexTTS-2 (64.34% to 77.29%) and 6.92 points for CosyVoice2 (70.90% to 77.82%) on IEMOCAP, described as relative gains of 20.1% and 9.8%.
- All four objective emotion metrics improve: The abstract states EmoRES outperforms CoCoEmo across all four objective emotion metrics on both IndexTTS-2 and CosyVoice2 on IEMOCAP. TEP moves from 36.22% to 41.11% and E-SIM from 65.32% to 67.67% for IndexTTS-2; TEP from 37.44% to 40.12% and E-SIM from 63.43% to 65.34% for CosyVoice2.
- Subjective emotion labeling: Dom-hit rises from 54.72% to 73.89% for IndexTTS-2 and from 58.18% to 77.88% for CosyVoice2 relative to CoCoEmo; Fidelity rises from 60.05% to 66.32% and from 50.80% to 59.60%. Without steering, native conditioning is weak, especially for IndexTTS-2 (Dom-hit 17.78%, Fidelity 25.04%).
- Naturalness preference: Listeners prefer EmoRES for naturalness with preference scores of 63.80% (IndexTTS-2, 95% CI [61.70, 65.91]) and 60.32% (CosyVoice2, 95% CI [58.08, 62.56]), both significantly above the 50% indifference point. The abstract reports listeners prefer EmoRES in up to 63.8% of pairwise comparisons and a relative fidelity improvement up to 17.3%.
- Shared component alone is not enough: Shared-only steering (λ_r = 0) improves E-SIM and TEP over no steering but yields the lowest rank correlation and H-Rate among steered conditions: ρ falls to 5.30% for IndexTTS-2 and 7.20% for CosyVoice2. As a separate diagnostic, it reduces the recognizer's neutral posterior from 36.51% to 7.49% for IndexTTS-2.
- Residual component alone is also not enough: Residual-only steering (λ_c = 0, λ_r = 1) improves ρ and H-Rate relative to λ_c = λ_r = 1 but reduces E-SIM and TEP on both models.
- Rebalancing beats scaling: With steering magnitude and vector norm held fixed, retaining the shared component and raising residual strength to λ_r = 3 gives the best emotion control; gains beyond λ_r = 3 are smaller, suggesting diminishing returns. CoCoEmo largely saturates beyond α = 3 on emotion metrics.
- Embedding structure: On the single-emotion CREMA-D subset (n = 486), probe accuracy is 56.9% for no-steer, 57.6% for shared-only, 81.4% for residual-only, and 86.9% for EmoRES; cosine silhouette score is 0.379 for residual-only and 0.460 for EmoRES, the highest reported.
- Speaker and intelligibility effects: EmoRES improves speaker similarity in every comparison in Table 1 and reduces WER under most settings, but on CREMA-D with CosyVoice2 WER slightly increases from 1.21% to 1.30%.
Methodology in Plain English
The method has two stages. First, in an offline, one-time extraction step, the researchers take pairs of emotional and neutral recordings that share the same speaker and the same words, from ESD, CREMA-D, and RAVDESS, and average the internal activations of the TTS model. The emotional mean minus the neutral mean gives an emotion vector for each emotion (angry, happy, sad, surprise, with neutral as the reference). Every pair is filtered by a speech-emotion-recognition gate and a signal-quality gate, and pairs are discarded whole if either side fails.
Second, at inference time the chosen vector is added to the activation at the same layer and site it was extracted from, and the hidden state's original magnitude is restored afterward by rescaling, so only the direction changes. Two frozen backbones are used: IndexTTS-2 (embedding-conditioned, layers 1, 6, and 8 of its semantic language model) and CosyVoice2 (instruction-conditioned, layers 14 and 17 of its text-to-token language model). These layer sets are the top-K most emotion-separable layers published by CoCoEmo.
The core idea is a decomposition. The authors compute the centroid of all emotional activation means and split each emotion vector into a shared part (neutral to centroid, identical for every request) and a residual part (centroid to the specific emotion, which distinguishes one emotion from another). EmoRES assigns separate coefficients λ_c and λ_r to these two parts instead of one global strength, and sets the residual coefficient higher (λ_r = 3 in the main experiments, λ_c = 1). To keep comparisons fair, every steered condition is rescaled to the same vector norm, so differences come from direction and component balance, not from injecting a larger vector. Tuning uses CREMA-D (annotated with angry, happy, sad mixtures) as the development set; IEMOCAP (annotated with angry, happy, surprise, sad) is held out entirely as the out-of-distribution test set. Baselines are no steering, CoCoEmo, and random isotropic Gaussian steering. Objective metrics are target emotion probability (TEP), weighted anchored emotion similarity (E-SIM), speaker similarity (S-SIM), word error rate (WER), Spearman rank correlation (ρ), and hit rate (H-Rate), with emotion2vec+ large as the speech emotion recognizer. Human evaluation collects emotion labels and blind pairwise naturalness judgments, with at least 9 annotations per sample.
Why This Matters
Impact on research. The paper reframes emotion steering as a compositional problem rather than a single-direction injection problem, and it shows that an existing method sits inside a more general family. Because it requires no retraining, no new learned subspaces, and no target-emotion reference recording, it is a low-cost way to strengthen controllability on top of existing zero-shot TTS systems. The finding that the shared component mainly moves speech away from neutral, while the residual carries which emotion is expressed, is a testable structural claim about the geometry of emotion representations in speech language models.
Real-world applications:
- Conversational agents that need to convey a specific affect reliably.
- Narration and audiobook or content-creation systems where expressive control matters.
- Accessibility tools that communicate affect through synthesized speech.
- Dubbing, where a requested emotion must be conveyed without compromising naturalness, intelligibility, or speaker identity.
Industry relevance. Both backbones are modern zero-shot TTS systems with different conditioning interfaces (emotion embeddings versus natural-language instructions), so the method is presented as applying across interface designs. The code is released at a public repository, and the approach adds inference-time control without touching the frozen backbone's weights, which matters for deployment where retraining or fine-tuning is impractical.
Future Directions
- Extend EmoRES to token-level or segment-level control for emotions that vary within a single utterance, as stated in the conclusion.
- Test whether the shared-residual decomposition generalizes to additional emotions, languages, speech corpora, and TTS architectures.
- Examine whether the rebalancing benefit persists with other speech emotion recognizers, given the paper includes an appendix on evaluating with different recognizers.
- Explore control settings where native emotion conditioning is not used at all, addressed in the paper's appendix on steering without native emotion conditioning.
Target Audience
Researchers and engineers working on expressive or emotional speech synthesis, on activation steering and interpretability of speech language models, and on inference-time control of frozen generative models. It is also relevant to practitioners deploying zero-shot TTS who need stronger emotion adherence without retraining. Readers without background in TTS architectures or representation steering will find the method harder to follow, since it assumes familiarity with hidden activations, layers, and norm-preserving interventions.
Authors’ abstract
Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.