RECAST:用 RouterLM 学习主动计算证据,不再依赖固定检索
RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
这篇论文把 RAG 从'检索'升级成'主动算证据',RouterLM 用 Qwen3.5-9B 就打赢了 Gemini 3.5 Flash,做 RAG 的可以看看思路。
arXiv 论文 RECAST 将 RAG 中的证据构建改为序贯决策过程,由 RouterLM 迭代选择检索或计算操作,再交给冻结的 CompilerLM 生成可执行代码、AnswerLM 作答。RouterLM 采用 SFT 加 GRPO 训练,在 6 个基准家族上平均成功率达 75.6%,比最强基线高 15.9%。训练后的 Qwen3.5-9B RouterLM 超过免训练的 Gemini 3.5 Flash RouterLM 5.0%。在 3 个 held-out 基准上,RECAST 平均领先最强基线 15.0%,显示零样本泛化能力。
RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
Large language models are increasingly applied to tasks grounded in long, heterogeneous information sources. Conventional Retrieval-Augmented Generation (RAG) relies on fixed similarity-based retrieval, while agentic variants adapt queries and tool use but remain largely retrieval-centric. However, in many tasks, the evidence required for a solution is not explicitly present in any single source item. Instead, it must be derived through filtering, aggregation, or computation across multiple source items. In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved. A lightweight RouterLM iteratively selects and formulates primitive operations or specifies customized operations for a frozen CompilerLM to translate into executable code. Once it judges the evidence sufficient, RouterLM passes the accepted evidence to a frozen AnswerLM to produce the final solution. We train RouterLM with supervised fine-tuning (SFT) followed by group relative policy optimization (GRPO). Across six heterogeneous benchmark families, RECAST achieves a mean success rate of 75.6%, outperforming the strongest large-model baseline by 15.9%. Moreover, training enables the Qwen3.5-9B RouterLM to outperform a training-free Gemini 3.5 Flash RouterLM by 5.0%. On three held-out benchmarks, RECAST improves over the strongest baseline by 15.0% on average, demonstrating strong zero-shot generalization across tasks and heterogeneous source representations.