Cognition 评测:GPT-6 Sol 读取上下文比 GPT-5.6 Sol 少约 17%
In our evaluations, we see that GPT-6 Sol reads about 17% less context than GPT-5.6 Sol to reach the...
Cognition 跑了 GPT-6 两个型号的编程实测:Sol 省了 17% 上下文还少留烂尾测试,Luna 会写真回归测试了,写代码的可以看看。
Cognition 发布了对 GPT-6 系列模型的编程评测结果。数据显示,GPT-6 Sol 达到与 GPT-5.6 Sol 相同分数所需的上下文读取量少约 17%,且最后一次编辑后很少在仓库中留下失败测试。GPT-6 Luna 的提升集中在低、中推理档位,此时它会编写真实的回归测试而非基于 mock 的测试。
In our evaluations, we see that GPT-6 Sol reads about 17% less context than GPT-5.6 Sol to reach the...
In our evaluations, we see that GPT-6 Sol reads about 17% less context than GPT-5.6 Sol to reach the same score, and rarely leaves a repo with failing tests after its last edit. GPT-6 Luna’s biggest gains are at low and medium reasoning effort, where it writes real regression tests instead of mock-based ones. Read more: devin.ai/blog/gpt-6-sol… 💬 1 🔄 0 ❤️ 7 👀 1017 📊 2 ⚡