技巧精选73°

从后排放映:印度语数据、视频 token、PDF 解析与视觉 OCR 干货合集

From the back row: - Well under one percent of the web crawl frontier models learn from is in an In...

精选理由

一次看完五个视觉与 OCR 实战案例:Sarvam 读了 13 万亿 token,Reducto 解析 PDF 能提分省 token,Noyan 花几美元训检测器,做法都够具体。

Sarvam 的文档模型在看到任何图像前先读了 13 万亿 token 文本,上线四个月已数字化 3500 万页。Perceptron AI 针对一小时视频约百万视觉 token、监督信号仅约 2% 的问题,让模型自己决定读取哪些 token。Reducto 发现把 PDF 页面先解析再喂给前沿模型,可将约 30% 的 PDF 决策基准成绩提升并减少推理 token。Merve Noyan 用大开源模型标注图像、两个小模型筛选框、再训练检测器,成本仅 3-4 美元。Meta 的三智能体审核管线通过压缩相似帧、缓存热门视频判定结果来过滤大部分视频。

图片来源 · AI Engineer
原文 · AI Engineer

From the back row: - Well under one percent of the web crawl frontier models learn from is in an In...

From the back row: - Well under one percent of the web crawl frontier models learn from is in an Indian language. Sarvam's document model read thirteen trillion tokens of text before it saw a pixel, and four months after launch it is digitizing 35 million pages. - Feed a model an hour of video and roughly a million visual tokens go in. The only ground truth is a transcript or a few labeled frames, so it learns from about two percent of them. Perceptron AI's fix is a model that decides which tokens to read. - A frontier model scores about thirty percent on a data lab's benchmark of decisions made from PDFs. Reducto found that handing models a parsed version of the page lifted their scores and cut their reasoning tokens. - Merve Noyan's pipeline labels images with a big open model, has two smaller models judge the boxes, and trains a detector on the survivors, for three or four dollars. Her coding agent made extra training images by flipping traffic signs left to right, until told not to. - Meta keeps most videos out of its three agent review pipeline. It compresses similar frames, caches verdicts on viral videos, and lets creators with a strong record skip it on metadata alone. Vision & OCR playlist: youtube.com/watch?v=RQi7x-… 💬 1 🔄 0 ❤️ 2 👀 390 📊 2 ⚡