NVIDIA搞了个Spatial-IQ基准,把数箱子拆成九个子任务,Qwen2.5-VL练完后准确率从2.9%飙到62.6%,但离人类82.1%还远。
NVIDIA Research发布空间推理诊断基准Spatial-IQ,将3D物体计数拆分为九项感知与认知子任务。人类在该基准上计数准确率为82.1%,而最好的现成多模态模型仅达17.7%。针对子任务训练后,Qwen2.5-VL-32B的计数准确率从2.9%提升到62.6%。该基准可帮助研究者定位空间推理失败点并验证能力改进。
Humans count the boxes in this image, hidden ones included, with 82.1% accuracy. The best off-the-sh...
Humans count the boxes in this image, hidden ones included, with 82.1% accuracy. The best off-the-shelf multimodal model manages 17.7%. Spatial-IQ is a diagnostic benchmark from NVIDIA Research that breaks 3D object counting into nine perceptual and cognitive sub-tasks, from counting columns to inferring the blocks that must be underneath to hold the structure up, and scores each one separately. Training on those sub-tasks lifted Qwen2.5-VL-32B object-counting accuracy from 2.9% to 62.6%. For researchers developing multimodal reasoning systems, this creates a practical loop: identify where spatial reasoning fails, target the missing capability, and verify that improvement reflects composition, not just a better final score. 💬 6 🔄 2 ❤️ 23 👀 3023 📊 9 ⚡
- Hunyuan08:32原文