基于AI的音效生成:跨输入模态生成模型综述

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

精选理由

想了解AI怎么根据文字或画面生成音效?这篇综述把近五年的模型都梳理了一遍,还点出了哪些任务还没做好,适合做声音设计的人参考。

AI 摘要

这篇综述从Google Scholar、IEEE Xplore和ACM数字图书馆筛选了30篇同行评审文章,系统梳理了文本、视觉、音频和多模态输入对生成音效质量的影响。过去五年中,多个生成模型在语义对齐和时间连贯性上达到SOTA水平。但复杂多事件场景的时序同步仍是难点,客观指标与人类感知存在差距,可控性和生成多样性之间也有权衡。

原文 · arXiv cs.AI

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial intelligence (AI)-driven audio generative models are rapidly growing in popularity and have the potential to transform the way sound is synthesized and used across various applications. In response to this growing momentum, this chapter reviews and analyzes recent AI-based generative models for sound effect synthesis, with a focus on how different input modalities (text, visual, audio, and multimodal) affect the quality, controllability, and contextual relevance of the generated audio. It examines 30 peer-reviewed articles sourced from Google Scholar, IEEE Xplore, and the ACM Digital Library, exploring the evolution of AI generative models over the past five years. The results show that multiple models achieved state-of-the-art performance, producing high-fidelity, semantically aligned, and increasingly temporally coherent sound effects across tasks. However, despite these advances, the review identifies persistent challenges, including limitations in temporal synchronization for complex multi-event scenarios, gaps between objective metrics and human perception, and trade-offs between controllability and generative diversity. Overall, the chapter highlights that AI-driven sound effect generation is progressing toward more adaptive, scalable, and context-aware systems, offering significant implications for future sound design workflows and interactive media applications.