Ornith-1.0 开源了从9B到397B的编程模型,在SWE-Bench等基准上超越Claude Opus 4.7,还能自己优化任务框架。
Ornith-1.0 系列开源模型发布,专门用于 agentic coding,参数从9B Dense到397B MoE全覆盖。在 Terminal-Bench 2.1 上得分77.5,SWE-Bench verified 82.4,NL2Repo 48.2。397B MoE模型在多个基准上超过 Claude Opus 4.7。模型采用自改进训练策略,利用强化学习同时生成解决方案和 task-specific scaffold。基于 gemma4 和 qwen3.5 后训练,MIT 许可开源。
这个也算是模型的大新闻了: Ornith-1.0 系列开源 模型发布了,主打编程,模型不仅生成代码解决方案,还能自己生成和优化 task-specific scaffolding 能力。 从 9B ...
这个也算是模型的大新闻了: Ornith-1.0 系列开源 模型发布了,主打编程,模型不仅生成代码解决方案,还能自己生成和优化 task-specific scaffolding 能力。 从 9B 到 397B 全覆盖,在 Terminal-Bench 2.1 和 SWE-Bench 等 agentic coding 基准上达到开源 SOTA,397B 模型超过 Claude Opus 4.7,在多个 benchmark 上表现突出,9B 小模型表现也很不错,适合边缘部署。 开源 agentic coding 又多了一个可玩的选择。 Ornith @ornith_ Aloha! 🌺 Meet Ornith-1.0, a family of open-source LLMs specialized for agentic coding. Ornith-1.0 spans the full parameter sizes including 9B Dense, 31B Dense, 35B MoE, and 397B MoE. It achieves state-of-the-art performance among open-source models of comparable size on coding benchmarks including: ✅Terminal-Bench 2.1(77.5) ✅SWE-Bench(82.4 on verified, 62.2 on pro, 78.9 on Multilingual) ✅NL2Repo(48.2) ✅SWE Atlas(41.2 on QnA, 42.6 RF, 39.1 TW) ✅ClawEval(77.1) Post-trained on top of gemma4 and qwen3.5, Ornith-1.0 employs a novel self-improving training strategy in which reinforcement learning is used to generate not only solution rollouts, but also the task-specific scaffolds that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model generate higher-quality solutions in agentic coding.😎 All models are released under the MIT license, enabling full commercial and research use. 📖Tech Blo deep-reinforce.com/ornith_1_0.html WFn 🤗Huggingfa huggingface.co/collections/de… eBtM 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 170 ⚡