论文09:28InterFLOPBench: LLMs浮点错误分类F1超0.88这篇论文搞了个InterFLOPBench,专门测LLM找C语言浮点bug。Qwen 3和Gemini 2.5等模型表现最强,F1超0.88。写代码的必看!#InterFLOPBench#LLM#浮点错误#基准测试aarXiv: DeepSeek@Lisa Taldir 等 5 人原文稍后读已读值得跟进有用关注 InterFLOPBench