想试最强编程+安全模型?GLM-5.3刚发布,SWE-Marathon和CyberGym都拿了第一,看它怎么碾压对手。
智谱AI首席科学家唐杰在X上回应‘sooooooon’数小时后,GLM-5.3正式上线。该模型专注编程与安全,将SWE-Marathon成绩翻倍至42.5,Terminal Bench 3.0提升至28.3,是原来的5倍。在CyberGym基准上,GLM-5.3以84.5分居所有模型之首。模型发布距用户调侃唐杰仅三天。
GLM-5.3 Arrives Hours After Tang Jie's 'sooooooon' — Zhipu's Coding and Security Model Doubles SWE-Marathon, Tops CyberGym
Three days after a user teased Zhipu AI chief scientist Tang Jie on X about GLM-5.3, he replied 'sooooooon', and hours later the model was live. Focused solely on coding and security, GLM-5.3 more than doubled SWE-Marathon to 42.5, quintupled Terminal Bench 3.0 to 28.3, and scored 84.5 on CyberGym, first among all models.