这篇论文提出了SIEVE,用智能体主动找几处关键证据就能判断视频真假,不用看完整个视频,效果更好还更透明。
该论文提出SIEVE框架,用于多模态视频虚假信息检测。SIEVE通过一个证据搜寻智能体主动从视频中获取稀疏但关键的线索,构建紧凑的证据包。该智能体通过监督式证据搜寻轨迹和证据感知强化学习进行训练,鼓励获取信息性证据并抑制无效交互。在多个视频虚假信息检测基准上,SIEVE持续优于基线方法,并支持仅用紧凑证据包完成可靠验证。其证据获取过程提供了可检查的轨迹,增强了检测的透明度和可解释性。
Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection
Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its associated content are processed and judged in a single pass. However, real-world misinformation often exhibits a sparse and compositional evidence structure: a reliable decision may depend on only a few coupled clues, while most video content contributes limited additional information. Exhaustive multimodal reasoning may therefore introduce substantial redundancy and obscure decisive evidence. This motivates decoupling evidence acquisition from verification: first identifying sparse, decision-relevant clues and then judging veracity based on the acquired evidence. Accordingly, we propose SIEVE, a framework for Sparse Interactive Evidence Verification via Extraction in multimodal video misinformation detection. An evidence-seeking agent actively explores the available multimodal evidence and constructs a compact evidence package, which is then used by a verifier to determine veracity. The agent is trained with supervised evidence-seeking trajectories and an evidence-aware reinforcement learning objective that promotes informative evidence acquisition while discouraging unnecessary or invalid interactions. Experiments on multiple video misinformation benchmarks show that SIEVE consistently outperforms the evaluated baselines and supports reliable verification using compact evidence packages. Moreover, the resulting acquisition process provides an explicit and inspectable evidence trail, improving the transparency and groundedness of multimodal misinformation detection.