论文

JOVE:执行与验证联合调度,降低 LLM 任务图成本与延迟

JOVE: Joint Execution and Verification for Resource-Aware LLM Task Graphs

精选理由

把复杂任务拆成图后该派给哪个模型干?这篇论文给出了一套边执行边验证的在线调度方法,成本延迟降了 3 倍多。

论文提出 JOVE 框架,把复杂推理查询分解为有向无环任务图,在异构 LLM 之间做在线分配。系统通过异步付费验证反馈更新各 LLM 的质量估计,并用信息增益奖励把学习价值纳入分配决策。每条查询通过混合整数线性规划求解执行与验证决策,在长期预算和延迟约束下平衡开销。在四个推理基准上,JOVE 在保持有竞争力的准确率的同时,平均成本和延迟至少降低 3.17 倍,并证明了次线性的质量学习遗憾界。

原文 · arXiv cs.AI

JOVE: Joint Execution and Verification for Resource-Aware LLM Task Graphs

Complex reasoning queries can be decomposed into directed acyclic task graphs and distributed across heterogeneous LLMs, reducing latency through parallelism and enabling smaller models to solve complex tasks. In practice, however, the suitability of an LLM for a given subtask may be a priori unknown, and execution alone does not reveal output correctness. We propose JOVE, an online framework that jointly assigns executor LLMs and selects intermediate outputs for paid verification. Verification runs asynchronously and is used to improve future allocations, so the system must balance spending on execution now against learning for later. We study how to optimize this trade-off under a long-term budget and a per-query latency constraint, with stochastic, initially unknown LLM service quality, invocation costs, and execution times. JOVE makes execution and verification decisions by solving a sequence of per-query mixed-integer linear programs. Online learning updates task-dependent estimates of LLM quality based on verification feedback, while an information-gain bonus incorporates the value of learning into allocation decisions. Under a natural set of assumptions, we establish sublinear quality-learning regret for JOVE. Across four reasoning benchmarks, JOVE achieves competitive accuracy against standard inference baselines while reducing average cost and latency by at least 3.17 times.