AI模型精选

webAI开源3B推理模型TwiL-LM3,4项基准超GPT-OSS-120B

webAI Intelligence Lab 开源了其首个形式化推理模型 TwiL-LM3,参数量仅 3B 模型在 5 项形式化推理基准测试中,有 4 项超过了 OpenAI 的 GPT-OSS-1...

精选理由

3B 的小模型在推理上干翻了 120B 的大家伙,还能跑在树莓派上,开源了直接就能玩。

AI 摘要

webAI Intelligence Lab 开源了首个形式化推理模型 TwiL-LM3,参数量仅 3B。在 5 项形式化推理基准中,TwiL-LM3 有 4 项超过 OpenAI 的 GPT-OSS-120B,参数量约为后者的 1/40,推理速度快 2.6 倍。该模型基于 webAI 自有的验证数据集训练,而非抓取互联网数据,可在树莓派或 iPhone 等消费级硬件上运行。

原文 · shao__meng

webAI Intelligence Lab 开源了其首个形式化推理模型 TwiL-LM3,参数量仅 3B 模型在 5 项形式化推理基准测试中,有 4 项超过了 OpenAI 的 GPT-OSS-1...

webAI Intelligence Lab 开源了其首个形式化推理模型 TwiL-LM3,参数量仅 3B 模型在 5 项形式化推理基准测试中,有 4 项超过了 OpenAI 的 GPT-OSS-120B(120B 参数),参数量约为后者的 1/40,推理速度快 2.6 倍。 开源地址: huggingface.co/webAI-Official… David Stout @Davidstout Today, we’re excited to open-source TwiL-LM3, the first formal reasoning model from the webAI Intelligence Lab. At just 3 billion parameters, TwiL-LM3 outperforms OpenAI’s GPT-OSS-120B on 4 of 5 formal reasoning benchmarks while running efficiently on consumer hardware. That’s 40× fewer parameters, 2.6× faster inference, and state-of-the-art performance in the reasoning tasks that power reliable tool calling, code generation, structured outputs, and AI agents. TwiL-LM3 was trained using webAI’s proprietary reasoning pipeline on webAI-owned, verified datasets—not scraped internet data. We believe better reasoning comes from better training pipelines and higher-quality data, not simply larger models. Our approach demonstrates that efficient models can rival—and in many cases surpass—models dozens of times their size. Designed for the edge, TwiL-LM3 runs on hardware people already own—from a Raspberry Pi to an iPhone—bringing advanced reasoning to millions of devices without relying on the cloud. This is our first open-source release from the webAI Intelligence Lab, and it’s only the beginning. Proudly built in Austin, Texas. Article: webai.com/blog/webai-rel… 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 1 👀 206 📊 1 ⚡