LingBot-Video开源了!MoE架构30B参数推理只要3B,物理推理基准超Wan2.6。做具身智能的赶紧试试。
LingBot-Video是首个基于MoE的具身智能视频基础模型,总参数量30B,推理时仅激活3B。模型在70K小时具身数据上微调,并集成大规模互联网视频预训练。在RBench基准上,LingBot-Video超越了Wan2.6、Seedance 1.5 Pro和Cosmos3 Super。其设计优化物理推理而非视频画质,通过稀疏激活大幅降低长视频推理成本。
Massive open-source release! LingBot-Video isn't about video quality; it optimizes for physical rea...
Massive open-source release! LingBot-Video isn't about video quality; it optimizes for physical reasoning, and it runs sparse: only 3B of 30B params are active at inference. MoE is finally showing up in long-context video, where inference cost matters a lot. Worth a look. Robbyant @robbyant_brain Today we open-source LingBot-Video — the first MoE-based video foundation model built for embodied intelligence. 🔹30B params, only 3B active at inference. 🔹Augmented with 70K hours of embodied data on top of large-scale internet video pretraining. 🔹Already outperforming Wan2.6, Seedance 1.5 Pro, and Cosmos3 Super on RBench. 🧵👇 Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 2 ❤️ 14 👀 2040 📊 3 ⚡