论文多源确认

Faynt:任天堂格斗游戏AI模型发布

Faynt: Scaling and Optimizing Policies for Competitive Melee

精选理由

Faynt是首个能控制全部26个任天堂格斗游戏角色的AI模型,胜率超98%,还开源了代码和锦标赛平台。

Faynt是专为《任天堂明星大乱斗》开发的Transformer模型家族,包含1000万和7500万参数两个版本。该模型可控制全部26个角色,在244场同角色比赛中赢得240场,胜率达98.4%。经过强化学习后,1000万参数模型在零延迟条件下击败了Slippi-AI模型的所有68场比赛。模型在NVIDIA T4上的推理延迟分别为5.2毫秒和8.7毫秒。

原文 · arXiv cs.LG

Faynt: Scaling and Optimizing Policies for Competitive Melee

We introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee, each controlling all 26 characters with a single checkpoint. After reinforcement learning (RL), the 10M wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases on their supported rosters, with a winning record against every release. These opponents retain 21- or 24-frame action delays; Faynt uses no added delay, and we have not isolated the effect of this difference. In a separate evaluation against a privately supplied zero-delay Slippi-AI model, the 10M wins all 68 games across two conditioning settings. We study architecture, optimization, scaling, and hyperparameter transfer to guide pretraining on approximately 840,000 human replays. Post-training combines rank- and outcome-based curricula, 75M-to-10M distillation, and RL restricted to Fox mirror matches. On the initial 152-game benchmark, the supervised 10M wins 69.7% of games, compared with 45.4% for the pretrained 75M, despite higher overall held-out controller-prediction loss. The weighted validation loss used for supervised checkpoint selection agrees with the win-rate ordering of all four pretrained and supervised policies. After supervised post-training, both models take less damage per minute, build larger early leads, and win more often after losing the first life. Optimized inference on recorded game states averages 5.2 ms per decision for the 10M and 8.7 ms for the 75M on an NVIDIA T4, excluding emulator execution and communication. We open-source the weights, both benchmark suites, and a platform for automated model tournaments.