Gemini 3.6 Flash发布:DeepSWE提升至49%,成本降低52%

I had early access to Gemini 3.6 Flash and found i…

精选理由

想要又快又便宜的模型?Gemini 3.6 Flash在DeepSWE上冲到49%,成本砍半,跑agent任务很合适。

AI 摘要

Google发布Gemini 3.6 Flash,定位高效工作模型。在DeepSWE基准上得分从37%提升至49%,每任务成本降低52%,输出token减少65%。AI Index得分保持在50,推理成本降低17%,输出速度达280 tokens/s。该模型与GPT-5.6 Luna(51分)和Claude Sonnet 5竞争,但更快更便宜。Google优化了高吞吐量代理工作流的部署效率。

原文 · @koltregaskes

I had early access to Gemini 3.6 Flash and found i…

I had early access to Gemini 3.6 Flash and found it to be very fast, but I did have problems with it. I reported technical issues with the visualisations as well as the usual context loss I always have with Gemini models.

But first though let's make it clear: Flash is designed for speed hence the name Flash. Comparing it to GPT Sol, Meta Spark, Claude Fable, and the other top models is just stupidity. Let's put this into perspective.

Gemini 3.6 Flash launched yesterday. It's Google's efficient workhorse - faster, cheaper, better at coding and knowledge work.

This *isn't* a frontier model, nor is it designed to be. It sits in the mid-tier efficient category. But the efficiency gains over 3.5 Flash are big.

On DeepSWE it jumped from 37% to 49% while cutting cost per task by 52% and output tokens by 65%. On the Artificial Analysis Intelligence Index it held at 50 but got 17% cheaper and runs at 280 tokens/sec.

Like-for-like it's right alongside GPT-5.6 Luna (51) and competitive with Claude Sonnet 5. If it's competing with Sonnet 5 while being faster and cheaper, that's solid for a Flash-class model.

The flat Intelligence Index score while improving speed and cost means Google optimised for deployment rather than benchmarks. For high-volume agentic workflows where you need fast, cheap, reliable performance, that makes sense.

Stats below.