论文

MOBA-VL模型实现电竞实时精准解说

MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary

精选理由

MOBA-VL用游戏遥测数据训练,能准确捕捉击杀和目标等关键事件,比现有模型更精准解说电竞比赛。

研究人员推出MOBA-VL模型,这是一个90亿参数的视觉语言模型,专为多人在线竞技场(MOBA)游戏实时解说设计。该模型使用游戏遥测数据作为监督信号,通过事件定位多轮强化学习训练。在MOBACast-Bench基准测试中,MOBA-VL在完整比赛和片段测试中分别达到63.25和63.45的总体分数,优于StreamingVLM和DeepSeek-V4.1-Flash模型。事件定位训练将事件召回率从34.5提升至42.1。

原文 · arXiv: DeepSeek

MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary

Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives. To address this limitation, we use game telemetry, which records exactly when each event occurs, as a supervision signal. We introduce MOBA-VL, a 9B-parameter model trained on this signal with event-localized multi-turn reinforcement learning, which rewards the turns that describe each event. We also collect MOBACast, 860 professional matches (about 460 hours) across three MOBA games with word-level timestamped commentary, and MOBACast-Bench, a benchmark from held-out tournaments. On MOBACast-Bench, MOBA-VL achieves the highest Overall score on full matches (63.25 vs. 55.12 for StreamingVLM) and clips (63.45 vs. 56.22 for DeepSeek-V4.1-Flash). Event-localized credit also raises event recall from 34.5 to 42.1 over supervised fine-tuning. Code and data will be released, and demos are available on an anonymous project page at https://moba-vl.github.io.