Qwen 团队发布 Qwen3.8-Omni-Flash:跨文本音频视频的多模态智能体模型
Qwen 团队把智能体能力扩展到音视频了,Qwen3.8-Omni-Flash 能做视频剪辑和长视频翻译,还配了两个开源框架,做多模态应用可以看看。
Qwen Team 发布 Qwen3.8-Omni-Flash,一个原生多模态模型,针对跨文本、音频、视频的长程智能体任务训练,可完成视频剪辑和长音频/视频翻译。模型采用 Qwen3.8-Next 的稀疏 MoE 架构,上下文窗口达 100 万 token。通过共训练策略,模型在保持文本性能的同时把智能体能力迁移到音频和视频任务。同步开源两个框架:Qwen-MM-Plugins 为现有智能体框架增加音视频支持,Qwen-Live-Harness 提供带上下文与记忆管理的实时多模态交互、工具调用和子智能体委派。
The era of omni agents is upon us.
This is a great report by the Qwen Team on their omni-modal agents.
They present Qwen3.8-Omni-Flash, a natively multimodal model trained for long-horizon agent tasks across text, audio, and video, such as video editing and long-form audio and video translation.
It uses the sparse mixture-of-experts design of Qwen3.8-Next with a context window of one million tokens. A co-training strategy keeps text performance while carrying agent skills over to audio and video tasks.
Two open-source frameworks come with it. Qwen-MM-Plugins adds audio and video support to existing agent harnesses, and Qwen-Live-Harness handles real-time multimodal interaction with context and memory management, tool use, and sub-agent delegation.
Paper: https://t.co/PxFrmGby1t