GoogleDeepMind的Gemini 3.8 Flash在智能体评测中排名大幅提升,用户反馈明显更积极。
Gemini 3.8 Flash (High)在Agent Arena评测中较Gemini 3.7 Flash (High)提升5.94%,排名从第32位上升至第14位。该模型在"赞扬vs投诉"指标上提升14.78%,在"确认成功率"上提升11.12%,在"Bash恢复"能力上提升4.46%。Agent Arena评测使用网络、文件系统和终端工具衡量模型在真实世界长期任务中的表现。
Gemini 3.8 Flash (High) by @GoogleDeepMind significantly improved in Agent Arena over Gemini 3.7 Fla...
Gemini 3.8 Flash (High) by @GoogleDeepMind significantly improved in Agent Arena over Gemini 3.7 Flash (High), scoring +5.94% net improvement and ranking #14 overall—up from +0.84% and #32 overall for 3.7! Agent Arena measures net improvement relative to the average model across real-world, long-horizon tasks using web, filesystem, and terminal tools. We measure overall and dive deeper by five key signals: - Confirmed Success: explicit “yes, that worked” from users - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? - Bash Recovery: how it recovers from CLI errors - Tool Hallucination: does it call tools that don't exist Of the five key signals we track, Gemini 3.8 Flash (High) demonstrates a massive gain in Praise vs. Complaint: users are reacting to its output with noticeably more positive sentiment. There is notable improvement in Confirmed Success, and Bash Recovery. Steerability remains a relative weakness, though it has improved compared with 3.7. Gemini 3.8 Flash (High) vs. Gemini 3.7 Flash (High): - Praise vs. Complaint: +14.78% vs -1.60% - Confirmed Success: +11.12% vs +9.80% - Bash Recovery: +4.46% vs +0.61% - Steerability: -1.43% vs -5.38% - Tool Hallucination: Scores were nearly identical, with neither model showing issues Big congrats to the @GoogleDeepMind team on this release! Google DeepMind @GoogleDeepMind Two new Gemini models are here to help scale your AI agents and secure code: 🔘 3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. 🔘 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level vulnerability detection and automated patching. 🔗 View Quoted Tweet 💬 10 🔄 1 ❤️ 105 👀 8989 📊 15 ⚡