EcoFrame不用微调,靠熵和注意力决定该看哪些帧,在长视频任务上比BOLT准还快1.85倍,代码已开源。
EcoFrame是一个无需训练的视觉证据调度框架,让长视频理解模型按需调整帧预算。它用熵门控判断当前证据是否足够,并用注意力引导在信息密集区域补充候选帧。在Video-MME、LongVideoBench和MLVU上,EcoFrame在Qwen2.5-VL上平均准确率达64.4,超过BOLT的63.5。相比AKS和BOLT,EcoFrame获得1.85倍加速;相比agent方法A.I.R.,推理加速最高达13.5倍且精度相当。
When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based methods rely on static one-shot selection with fixed frame budgets and candidate pools, while agent-based schedulers achieve adaptivity through costly multi-round reasoning and interactive search. We propose EcoFrame, a training-free framework for low-overhead query-adaptive visual evidence scheduling. EcoFrame leverages the VLM's inference feedback to determine when to increase the frame budget and where to search for additional candidate evidence. Specifically, entropy-gated budget scheduling uses output uncertainty to stop early when the current evidence is sufficient or progressively expand the frame budget otherwise. Meanwhile, attention-guided candidate proposal converts frame-level attention into a temporal prior, enabling dense local search in informative regions while preserving global coverage when attention is diffuse. Experiments on Video-MME, LongVideoBench, and MLVU demonstrate that EcoFrame achieves a better accuracy--efficiency trade-off across multiple VLM backbones. On Qwen2.5-VL, EcoFrame achieves an average accuracy of 64.4, surpassing BOLT at 63.5, while providing a $1.85\times$ speedup over AKS and BOLT. Compared with the agent-based A.I.R., EcoFrame maintains comparable accuracy with up to a $13.5\times$ inference speedup. Code will be available at https://github.com/AK-DREAM/EcoFrame.