AI Engineer World's Fair 2026 视觉与 OCR 议题上线
Live now: our Vision & OCR Track from AI Engineer World's Fair 2026. A model that counts 32 white s...
AI Engineer 把 2026 大会的视觉 OCR 场次全放出来了,LlamaIndex 和 Hugging Face 的人都来讲实战,做文档和视觉应用的很值得刷一遍。
AI Engineer 公开了 World's Fair 2026 的 Vision & OCR Track 全部演讲视频。议题包括 LlamaIndex 的 Jerry Liu 讲文档上下文层、Hugging Face 的 Merve Nooyan 谈用 Skills 部署视觉语言模型、Reducto 的文档智能流水线等。Sarvam 还带来从零训练 3B 状态空间视觉模型的分享。
Live now: our Vision & OCR Track from AI Engineer World's Fair 2026. A model that counts 32 white s...
Live now: our Vision & OCR Track from AI Engineer World's Fair 2026. A model that counts 32 white squares on part of a chessboard. A file format that stores a table as a pile of line segments. Ten turkeys on the roof of a Tesla. Thesis: the models can see. They are still learning to look. youtube.com/watch?v=RQi7x-… - Building the Document Context Layer for AI Agents: @jerryjliu0 , LlamaIndex - Skill issue: stop deploying vision language models, use them with Skills: @mervenoyann , Hugging Face - Modality Misalignment and Originality Attribution in Short-Form Video: Aditya Gautam, Meta - From Ingestion to Agents: How AI Teams Build on Document Intelligence: Adit Abraham, Reducto - The Best Models Still Reason Like Toddlers: @andrewdai , Elorian - You're Not Thinking Big Enough: Rebuilding Food Systems with AI Agents: @cbmenefee , Firecrawl - From VLM/VLA's to Embodied Agents: @ArmenAgha , Perceptron AI - From Scratch to SOTA: Training a 3B State-Space Vision Model: @fewshotlearner , Sarvam 💬 3 🔄 0 ❤️ 5 👀 861 📊 4 ⚡
- Jerry Liu09-21 23:46原文