MosiAI 开源了一个0.9B的小模型,能一次性转录90分钟的多说话人音频,边缘设备就能跑,很适合本地部署。
MosiAI 团队开源 MOSS-Transcribe-Diarize-0.9B,仅 0.9B 参数,基于 Whisper-Medium 编码器和 Qwen3-0.6B 风格解码器。模型支持 128K 上下文,可一次性处理约 90 分钟长音频。输出直接包含时间戳和说话人标签,无需分块或单独说话人分离。还支持热词提升,用于人名、领域术语和代码切换场景。SGLang 已提供 Day-0 支持,可立即运行。
🎉 Meet MOSS-Transcribe-Diarize-0.9B from the @Mosi…
🎉 Meet MOSS-Transcribe-Diarize-0.9B from the @MosiAI_Official team, a 0.9B open-source, end-to-end model for long-form multi-speaker ASR. Day-0 support is now live in SGLang!
1️⃣ Audio-in, structured text-out: timestamps + speaker labels, no chunking or separate diarization 2️⃣ Whisper-Medium encoder + Qwen3-0.6B-style decoder 3️⃣ 128K context, up to ~90 min audio in one pass 4️⃣ Hotword boosting for names, domain terms & code-switching 5️⃣ 0.9B params, edge/on-device ready
Cookbook: https://t.co/vuko42we4z Run it now with SGLang!
- vLLM07-09 11:22原文