RunningTab:用环境侧标签帮 LLM 智能体追踪工作区任务进度
RunningTab: Direct Workspace Interaction with Environment-Side Tabs
一篇解决智能体“读完文件就忘”的论文,让环境而不是模型来记账,避免报告漏掉已提取的数据,思路挺新颖。
LLM 智能体通过终端直接搜索和读取工作区文件(DWI),但上下文窗口记不住任务要求、已读文件和列出未打开的文件,可能出现提取了数据却在报告里遗漏的情况。RunningTab 框架由环境维护一个每任务标签,智能体登记需求,环境记录每次读取的文件摘录及来源,未打开的文件列为候选。智能体可以在需求旁看到最匹配的摘录和未读候选,逐一解决或注明原因搁置,任务未完成时无法通过 finish check。在三个基准和三个 LLM 上,RunningTab 持续超过纯 DWI 和将记录放在模型内的基线。
RunningTab: Direct Workspace Interaction with Environment-Side Tabs
Much knowledge work produces new deliverables from files a workspace already holds, and LLM agents are beginning to take such work over. Through direct corpus interaction, an agent can search and read any of those files from a terminal with no indexing, and producing a deliverable from many of them in this way is what we call direct workspace interaction (DWI). Reaching the files, however, is only half the task: nothing keeps track of what the task asks for, what has been read, and what was listed but never opened, all of which slip through the context window without leaving a trace, so an agent may extract a figure and still deliver a report without it. To address this, we present RunningTab, a framework that equips direct workspace interaction with an environment-side tab: a per-task record of what the task still owes, kept by the environment alongside the agent. Specifically, the agent adds its requirements, while the environment records every file read as an excerpt with its provenance and every listed but unopened file as a candidate; the agent can then see each requirement beside its best-matching excerpts and top unopened candidates, resolve it against matching content or set it aside with a reason, and, should it try to finish with requirements still open, receive them in a finish check. We validate RunningTab on three benchmarks with three LLMs, where it consistently outperforms plain DWI and baselines that keep the record in the model, while its tab usually holds the values a deliverable needs once seen.