微软发布7款MAI新模型,MAI-Thinking-1对标Sonnet 4.6

Super excited to announce seven new world-class MAI models today. They represent what we consider a ...

精选理由

微软一口气推出7款新模型,覆盖推理、编程、图像三大方向,MAI-Thinking-1在推理和编码上直接对标Claude Sonnet 4.6和Opus 4.6,做AI应用或企业定制化模型的团队值得关注——尤其是Frontier Tuning让企业用更低成本获得超越GPT-5.5的效果。

AI 摘要

微软CEO Mustafa Suleyman宣布推出7款全新MAI系列模型,包括文本基础模型MAI-Thinking-1、图像模型MAI-Image-2.5及高效编程模型MAI-Code-1-Flash。MAI-Thinking-1拥有350亿激活参数的MoE架构,256K上下文窗口,在AIME 2025上达到97%,SWE Bench Pro上53%,与Opus 4.6持平,且盲测中整体质量优于Sonnet 4.6。该模型针对微软自研MAIA 200芯片优化,性能每美元提升30%,每瓦性能提升1.4倍。MAI-Code-1-Flash仅5B参数,SWE Bench Pro达51%,成本更低。微软还推出Frontier Tuning服务,允许企业定制专属模型,早期案例中为McKinsey定制模型以10倍低成本超越GPT-5.5。

原文 · Mustafa Suleyman

Super excited to announce seven new world-class MAI models today. They represent what we consider a ...

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks. - It’s a 35B active parameter MoE with a 256K context window. Independent human raters on Surge prefer it for overall quality in blind side-by-sides versus Sonnet 4.6, and it’s achieved 97% on AIME 2025, the key measure of its general-purpose reasoning abilities. - It's at 53% on SWE Bench Pro, placing it right alongside Opus 4.6 on one of the toughest coding benchmarks. - And since we co-designed our models with our own silicon, MAI-Thinking-1 is optimized on our MAIA 200 chip. Benchmarking head-to-head against the GB200, we see 30% better performance per dollar as well as a 1.4x performance-per-watt gain when running our MAI models on the MAIA 200 end-to-end. Next is MAI-Image-2.5 and its Flash variant. Two super strong models now at #2 on the leaderboards, surpassing the score of Nano Banana 2 on image editing. Last for now is MAI-Code-1-Flash, our new inference efficient coding model, especially tuned for VS Code and GitHub Copilot CLI. - Code-1-Flash achieves 51% on SWE Bench Pro, despite having just 5B parameters, putting it closer to Haiku in size but cheaper in cost. All of this is the foundation for Microsoft Frontier Tuning. It lets you customize our models to create custom, company-specific agents that only you control. You can make our model, your model. Your data. Your agents. Your moat. Early adopters are already seeing a difference. When we tuned our models for McKinsey’s tasks, MAI delivered the highest win rate, outperforming GPT-5.5 on quality, while being 10x lower on cost. Also really excited to be collaborating with the amazing team at Mayo Clinic to jointly train a new frontier AI model for healthcare. Our announcements today mark another milestone on the road to humanist superintelligence. You can learn more and about our other new models in our latest blog: microsoft.ai/news/building-… 💬 118 🔄 371 ❤️ 2443 👀 452077 📊 530 ⚡