精选理由
腾讯发了ProLaViT,让多模态模型像人一样一步步想问题,不用外部工具就能解决视觉推理错误,ECCV 2026保底。
腾讯BAC团队提出ProLaViT框架,用于渐进式潜在视觉思维,使多模态LLM(如Qwen2.5-VL)无需外部视觉工具即可在潜在空间中逐步推理。该框架被ECCV 2026接收,在多个视觉推理基准(如MMMU、MathVista、CLEVR)上相比直接推理方法提升5-15%准确率。ProLaViT通过引入潜在思维令牌,让模型在内部推理步骤中分解复杂视觉问题。
原文 · Pandaily
Tencent ProLaViT Teaches Multimodal LLMs to Reason Step-by-Step in Latent Space, Ending Swallowing-Whole Visual Reasoning Failures
Tencent BAC team proposes ProLaViT framework for progressive latent visual thought, enabling multimodal LLMs to reason step-by-step in latent space without external vision tools, accepted at ECCV 2026.