谷歌突袭发布Gemini 3.7 Flash,AutomationBench拿到30%,比贵一倍的模型还强。做智能体自动化任务可以试试。
谷歌发布Gemini 3.7 Flash,在Zapier的AutomationBench基准上拿下30%得分,成为首个突破30%的模型。三周前Gemini 3.6 Flash得分19.8%,这次每任务成本仅增加3美分。在营销、财务、销售、支持四个领域表现领先,其中营销38%、财务37.5%。但在运营领域仍落后,Opus 5以50%保持领先。
If you are going to test/use Gemini 3.7 in @NousResearch Hermes or @openclaw. Please let me know how...
If you are going to test/use Gemini 3.7 in @NousResearch Hermes or @openclaw . Please let me know how it goes. Wade Foster @wadefoster Surprise drop today: Gemini 3.7 Flash. On AutomationBench, it beats models that cost twice as much. The progress here is insane. 3 weeks ago Gemini's 3.6 Flash scored 19.8%. Today, 3.7 is the first model to crack 30% on AutomationBench. At only 3 cents more per task. 𝗪𝗵𝗲𝗿𝗲 𝗶𝘁 𝘄𝗶𝗻𝘀: Marketing (38%), Finance (37.5%), Sales (28.2%), and Support (22%) Example: Confirm a project is done in the CRM, then run our standard label cleanup on its email threads. Archive the closed ones, leave restricted ones alone. 3.7 was a full pass in 22 steps. GPT-5.6 Sol failed the same task in 8, then hallucinated a summary. 𝗪𝗵𝗲𝗿𝗲 𝗶𝘁 𝗹𝗼𝘀𝗲𝘀: Operations. Opus 5 still runs that domain at 50%. Example: On a revenue-attribution task, 3.7 ran the math, updated the records, posted the summary, sent the escalation, and then never wrote the one required row on a second tracker. Came close, but still failed. How @Zapier 's AutomationBench works: we score every new model on 657 of the hardest workflows we run. Scoring is deterministic: either the right records got updated and the right messages got sent, or they didn't (no partial credit) @GeminiApp ’s 30% is a new record. See every zapier.com/benchmarks https://t.co/4ylnVysHFq 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 248 ⚡