用可执行真值逐帧审计视频模型计数,发现Gemini 3.6 Flash处理瞬时事件几乎失效,加帧只虚涨分数,值得看。
新论文提出trace-grounded参数化测试方法,在2,190个视频上审计事件计数能力。Gemini 3.6 Flash在80%可靠性阈值下最多可靠计数12个持续状态转换事件(0.5和1.0 Hz),对瞬时闪烁事件则无可靠计数区间。高计数高频率场景下仅0.2%的最终计数正确,模型只能恢复18.1%的真实事件。提高采样率将Bounce Ball准确率从19.6%提升至29.3%,但报告序列与真值一致率仅3.7%。
The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping
Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported events against executable ground truth. To bridge this gap, we introduce trace-grounded parametric profiling for event counting in three controlled video tasks: bouncing-ball wall contacts, visual blinks, and categorical state transitions. Across 2,190 videos, we vary event count N and frequency F while holding rendering fixed. Each video includes an executable event trace for capability-surface estimation and timestamp-level evaluation. Our results reveal a staged temporal failure. At an 80% reliability threshold, Gemini 3.6 Flash reliably counts persistent state transitions up to 12 events at 0.5 and 1.0 Hz, yet demonstrates no reliable positive-count region for transient blinking events. Thus, event representation dictates whether a model initially accesses evidence -- a limitation that compounds as count and frequency increase. In the high-count, high-frequency regime, only 0.2% of final counts are correct and the model recovers just 18.1% of true events. To test if visual access is the primary bottleneck, we increase sampling rate. Although this boosts Bounce Ball accuracy from 19.6% to 29.3%, the reported sequence agrees with ground truth only 3.7% of the time. Extra frames can therefore inflate final scores without producing faithful event recovery. Different prompting strategies yield similarly limited gains, and real-world video evaluations show the same concentration of success at low event counts. Ultimately, trace-grounded profiling shifts video evaluation from aggregate accuracy metrics to a detailed diagnostic of where temporal reasoning fails.
- 向阳乔木08-06 03:44原文