Poolside 又出了个厉害编码模型,Laguna S 2.1 在本地就能跑,多项基准比大它几十倍的模型还强。
Poolside AI 发布最新 agentic coding 模型 Laguna S 2.1。在 Terminal-Bench 2.1 上得分 70.2,与 5-25 倍大小的模型并列甚至领先。在 DeepSWE 基准上得分 40.4,超越部分超 1T 参数的开源模型。该模型可在单个 DGX Spark 或 Mac 上本地运行。
You should probably go give @poolsideai a follow on Hugging Face. These folks are on a roll, releas...
You should probably go give @poolsideai a follow on Hugging Face. These folks are on a roll, releasing one agentic coding models every month and the latest one Laguna S2.1 is quite possibly the best coding model you can run locally (single DGX spark or Mac) atm Poolside @poolsideai Laguna S 2.1 is, as far as we can measure, the most capable agentic coding model in its weight class. On Terminal-Bench 2.1 it scores 70.2, sitting beside models 5–25x its size and ahead of several of them. And on DeepSWE from @datacurve , the hardest long-horizon benchmark we ran, Laguna S 2.1 scores 40.4, outperforming some open models with more than 1T parameters. For every score we publish today, we're releasing the full trajectory of every trial in the final evaluation set at trajectories.poolside.ai 🔗 View Quoted Tweet 💬 0 🔄 1 ❤️ 3 👀 695 📊 1 ⚡