论文

GS-Codec:用 Gaussian Splatting 替代量化瓶颈的神经语音编码器

GS-Codec: A Gaussian-Splatting Bottleneck for Neural Audio Coding

精选理由

有团队把 3D 重建里的 Gaussian Splatting 搬进语音压缩,不用 codebook 也能做到和 EnCodec、DAC 差不多的效果,思路挺新鲜。

GS-Codec 是一种新神经语音编码器,将 3D 重建中的 Gaussian splatting 改造为一维信号分解,用 1D 高斯基元的加权和替代传统的 codebook 量化。训练时通过内层优化循环拟合基元参数,并训练 GS Predictor Net 在单次前向传播中直接回归这些参数。同一份 checkpoint 无需重训练,即可通过调整基元数量和每参数比特深度实现码率控制。在相近码率下,GS-Codec 在 SIM、STOI、UTMOS 三项指标上持平或超过 EnCodec 与 DAC,WER 语义表现相当。

原文 · arXiv cs.AI

GS-Codec: A Gaussian-Splatting Bottleneck for Neural Audio Coding

Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook variants, follows the VQ-VAE template by partitioning the encoder latent through a learned codebook or a fixed scalar grid. We ask whether this partition is necessary. We introduce GS-Codec, a neural speech codec whose bottleneck is a parametric signal decomposition rather than a quantizer. We adapt Gaussian splatting from 3D scene reconstruction to one-dimensional latents. An inner optimization loop fits each encoder segment as a weighted sum of 1D Gaussian primitives. The decoder then reconstructs the waveform from the rendered sum. To avoid the cost of this iterative inner loop at inference time, we additionally train a lightweight GS Predictor Net that regresses the primitive parameters in a single forward pass. The encoder and decoder are trained end-to-end through the inner loop, with no quantizer anywhere in the training pipeline: the bottleneck is the decomposition itself, and scalar quantization is applied only post-training to the fitted parameters. Rather than relying on discrete codebook stages for bitrate control, our representation exposes a fine-grained rate-quality tradeoff: a single trained checkpoint supports post-training bitrate control by varying the number of primitives and the per-parameter bit depth, with no retraining required. GS-Codec matches or exceeds well-established open-source codecs such as EnCodec and DAC on speaker similarity (SIM), intelligibility (STOI), and perceptual quality (UTMOS) at comparable bitrates, while achieving comparable semantic performance (WER). Code and audio samples are available at https://ronaluf.github.io/gs-codec/