Z.ai发布GLM-5.3-Flash,原生多模态MoE,上下文窗口更大,性价比高,值得尝试。
Z.ai推出GLM-5.3-Flash,GLM-5系列首个原生多模态模型,320B总参数/18B激活MoE,1,048,576令牌上下文窗口,MIT授权Hugging Face权重,API定价0.15美元/输入和0.50美元/输出。在Terminal-Bench 2.1上得分84.3,在DeepSWE v1.1上得分63.4,使用混合KDA线性加NoPE稀疏MLA注意力减少注意力计算约3倍,KV缓存4.4倍于GLM-5.3。
Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series — a 320B-total / 18B-active MoE with a 1,048,576-token context window, MIT-licensed weights on Hugging Face, and API pricing at $0.15/M input and $0.50/M output. It scores 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1, using hybrid KDA linear plus NoPE sparse MLA attention to cut attention compute ~3× and KV cache 4.4× versus GLM-5.3. The post Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context appeared first on MarkTechPost .