论文

SteerSpeech:用激活转向控制语音生成中的情绪

Steerspeech: Activation Steering For Emotion Control In Generated Speech

精选理由

给 TTS 加了个情绪旋钮,不用重新训练模型,注入转向向量就能调情绪强度,Qwen3-TTS 上效果最高提升 7 倍。

Qwen3-TTS 等预训练 TTS 模型难以在推理时精确控制情绪,提示词和参考音频只能提供粗粒度调节。SteerSpeech 通过向隐藏层注入转向向量来控制情绪,为每个目标情绪训练一个低秩变换,同时保持 TTS 主干冻结。为了让监督信号穿过离散语音 token,作者设计了带直通估计器的两段式生成-回放管线。在 Qwen3-TTS 上的评测显示,目标情绪得分达到基线的 1.08x-7.12x,高强度转向下说话人身份保留率为 1.43x-1.46x。

原文 · arXiv cs.LG

Steerspeech: Activation Steering For Emotion Control In Generated Speech

Pretrained text-to-speech (TTS) models can generate expressive speech, but reliable inference-time emotion control remains challenging: prompts and reference audio offer coarse, inconsistent control, whereas specialized conditioning and model adaptation require costly training. We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations. For each target emotion we train a lightweight low-rank transform, using a multi-expert objective that encourages monotonic emotion control while preserving speaker identity and linguistic content, constraining steering drift, and keeping the TTS backbone frozen. To optimize through discrete speech tokens, we introduce a two-pass generation-and-replay pipeline using a straight-through estimator to backpropagate expert supervision through sampled tokens. At inference, a target-emotion steering direction is optimized with its respective transform and injected into the base TTS model. Objective and subjective evaluations with Qwen3-TTS across seen, unseen, and accented speakers show stronger continuous emotion control with limited speaker and content degradation. SteerSpeech achieves 1.08x-7.12x baseline target-emotion scores and for a representative emotion subjectively, it receives 78.1%-96.8% intensity preference and 1.43x-1.46x speaker-identity preservation at high steering strengths.