PhysWave统一了自然语言和轨迹控制,通过物理约束解决了现有方法违反声学关系的问题。
PhysWave是一种物理引导的潜扩散模型,用于可控文本到一阶声学(FOA)音频生成。该模型通过可微分声学前验(球谐波方向一致性和反平方距离一致性)增强扩散训练,并构建了包含30万段FOA音频的数据集。实验结果表明,PhysWave在保持竞争力的音频质量的同时,能够生成空间一致性更好的FOA音频。
PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation
Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. However, existing text-to-FOA methods are largely data-driven and may produce audio that violates acoustic relations between source direction and distance. They also separate descriptive and parametric control, forcing users to trade usability for precision. In this paper, we present PhysWave, a physics-guided latent diffusion model for controllable text-to-FOA generation. PhysWave unifies natural-language and trajectory control through a shared waypoint-caption representation, and augments diffusion training with two differentiable acoustic priors: spherical-harmonic direction consistency and inverse-square distance consistency. To support dynamic spatial generation, we further construct a 300K-clip FOA dataset with diverse sound categories and source trajectories. Extensive results show that the proposed priors help PhysWave generate spatially consistent FOA audio while maintaining competitive audio quality. Further analyses show that these physics priors improve spatial consistency during training and can also be used as inference-time guidance for training-free spatial refinement.