论文

TokenCast 论文:预测 LLM Agent 执行中的 Token 消耗

TokenCast: Forecasting Token Consumption During LLM Agent Execution

精选理由

跑 Agent 最怕 token 预算失控,这篇论文的 TokenCast 能边执行边预测消耗,实测省 21.3% 的 token,做 agent 的都该看看。

LLM Agent 执行同一任务时,不同运行的 token 消耗可相差一个数量级,因为工具反馈和上下文膨胀使成本难以预估。TokenCast 为每个执行段学习可组合的成本表示,同时记录自身消耗与带来的上下文增长,拼接相邻段即可得到累计估算。在 SWE-bench Verified 等 4 个任务套件、6 个 agent 模型共 96 种组合上,相对最强基线的平均绝对误差降低 14.5%,每次运行预测耗时均值 32.8 ms,且不需要额外 LLM 调用。离线预算控制回放中,在与固定预算策略相同的轨迹完成率下平均少用 21.3% 的 token。

原文 · arXiv cs.LG

TokenCast: Forecasting Token Consumption During LLM Agent Execution

When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.