DungeonBench用D&D战斗测AI战术推理,分单场和连战两关,发现大模型会打单场却管不好资源,想测AI决策的可以看看。
DungeonBench是一个面向D&D战斗的战术推理基准,覆盖2014版SRD中大部分战斗相关内容,保留行动经济、生物特性、战场几何等机制。基准分Encounter和Day两个轨道:Encounter评估单场战斗的局部战术,Day通过持久生命值、法术位、消耗品和短休息将战斗串联,考验资源调度。同一决策流支持启发式控制器、语言模型策略、学习型排序器和强化学习智能体。评测显示前沿语言模型策略在直接对战中常获胜,但在链接战斗日上暴露资源预算、休息时机和规则纪律的不足。
DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat
Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once. We introduce DungeonBench, a benchmark for tactical reasoning in Dungeons & Dragons combat, built to cover the vast majority of combat-relevant 2014 System Reference Document content whose effects can be resolved by the simulator while retaining mechanics that simplified combat simulators often abstract away. At each step, DungeonBench exposes a complete tactical observation, a pending decision, and an indexed list of executable options spanning movement, attacks, spells, reactions, objectives, preparation, and scarce resources. The task is to value legal choices whose consequences depend on action economy, creature traits, battlefield geometry, timing windows, and future encounters. DungeonBench has two tracks: Encounter, which evaluates local tactical play in single fights, and Day, which links encounters through persistent hit points, spell slots, consumables, preparation, and short-rest timing, forcing policies to trade off immediate tactical advantage against future survivability. The same engine-generated decision stream supports heuristic controllers, language-model policies, learned option rankers, and masked-action reinforcement-learning agents. We evaluate frontier language-model policies on this shared decision stream. Results show that full tactical observations do not saturate the benchmark: frontier policies often win direct encounters, but linked encounter days expose failures in resource budgeting, rest timing, and rule-aware tactical discipline.