LACE:一种动态帧率编解码器的分层压缩方法
LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs
一个关于音频编解码器的新方法,能提高 TTS 的效率,比之前的动态帧率方法更好。
LACE 是一种针对动态帧率编解码器的分层压缩方法,通过在量化层独立应用压缩步骤,实现了层特定的分割边界。该方法在 LibriTTS 数据集上的实验表明,其重建任务上的速率质量权衡优于先前方法,并提升了 TTS 推理效率。
LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs
Neural audio codecs are a key component in speech language modeling. However, their high frame rates lead to long sequence lengths, increasing computational costs. Dynamic frame rate codecs mitigate this by reducing the effective frame rate using a compression step to merge multiple frames together. However, most prior methods either operate on single-codebook codecs or apply a single compression step before multi-layer quantization. This forces all quantization layers to share the same segmentation boundaries, despite the residual embeddings at different quantization layers exhibiting different rates of change over time. We propose LACE (Layer-Adaptive Codec Encoding), a dynamic frame rate codec that applies an independent compression step at each quantization layer, enabling layer-specific segmentation boundaries. To use LACE tokens in downstream text-to-speech (TTS), we further introduce union alignment and boundary anchor mechanisms to make durations consistent across layers while preserving compression benefits. Experiments on LibriTTS show that LACE offers a better rate-quality tradeoff than prior dynamic frame rate methods on the reconstruction task and improves TTS inference efficiency while maintaining competitive synthesis quality. Our code is released as part of the ESPnet3 codec recipe.