Twelve Labs 用三件套解决了传统视频处理漏掉动态和因果的问题,想做视频检索和推理的可以看看这套方案。
Twelve Labs 在 Vector Space Day SF 上指出了视频处理的三种失败模式:Wrong Context、Wrong Memory 和 Wrong Reasoning。他们推出的 Marengo 检索模型把视频变成可搜索内容。Pegasus 视频语言模型将检索到的片段转化为结构化答案。Jockey 智能体框架负责视频语料库的工作流和记忆层。底层由 Qdrant 处理存储与检索。
Most people treat video like a bag of frames or a transcript with timestamps. But you end up missing...
Most people treat video like a bag of frames or a transcript with timestamps. But you end up missing the most important parts of it: - Motion - Causality - Temporal progression - The relationship between what happens and why it matters At Vector Space Day SF, James Le from @twelve_labs walked through 3 failure modes that come from this: - Wrong Context - Wrong Memory - Wrong Reasoning And he introduced the stack built to fix it: → Marengo: a retrieval model that makes video searchable → Pegasus: a video language model that turns retrieved moments into structured answers → Jockey: an agentic framework and memory layer for video corpus workflows With Qdrant handling the storage and retrieval underneath it all. Full talk is live on our YouTube channel: youtube.com/watch?v=i8xZeK… 💬 0 🔄 1 ❤️ 1 👀 46 📊 1 ⚡