AI模型精选

SeedRealtime:字节跳动原生音视频全双工大模型

ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model

精选理由

SeedRealtime一个模型能看能听能说,实时双向对话不轮流,字节跳动Seed出品。

AI 摘要

字节跳动Seed团队发布SeedRealtime,一款原生音视频全双工大语言模型。该模型在统一架构中融合音频、视频与文本,可对连续多模态流进行实时交互,而非逐轮问答。官方宣称三项突破,包括联合音视频理解。Seed表示这是迈向全模态交互的一步。

图片来源 · marktechpost
原文 · marktechpost

ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model

ByteDance’s Seed team has introduced SeedRealtime, a native audio-visual full-duplex LLM. The model fuses audio, video and text in a single unified architecture. It interacts in real time over continuous multimodal streams, rather than one turn at a time. Seed positions it as a step toward omni-modal interaction, and claims three breakthroughs: joint audio-visual understanding, […] The post ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model appeared first on MarkTechPost .