面向LLM智能体工作流并行分支的直接潜在空间合成

Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows

精选理由

并行合成提速2.5-11倍

AI 摘要

Parallel-Synthesis框架使合成器直接消费并行工作线程的KV缓存,避免文本拼接冗余。它通过缓存映射器校准独立分支缓存,并微调合成适配器以支持非顺序缓存接口。在9个数据集(数学、科学问答、代码生成、GAIA、多智能体数据库诊断)上,7个超越或持平文本合成基线,首token延迟降低2.5-11倍。该工作为并行智能体分支的高效合成提供了新接口。

原文 · arXiv cs.AI

Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows

Large language models increasingly serve as execution engines for agentic systems, yet they still consume context through a sequential text interface. This creates a mismatch with modern structured agent workflows, in which independent branches explore subtasks, retrieve evidence, or generate candidate solutions before a final synthesis step. Existing systems typically merge these branches by concatenating their textual outputs, which discards the parallel structure and incurs redundant prefill computation. In this work, we introduce Parallel-Synthesis, a plug-and-play framework that enables a synthesizer to directly consume the KV caches produced by parallel worker agents. Parallel-Synthesis combines a cache mapper that calibrates independently generated branch caches with a fine-tuned synthesizer adapter that enables generation from this non-sequential cache interface. We train Parallel-Synthesis using data that exposes the synthesizer to parallel cache contexts, teaches aggregation across cached branches, and distills reasoning behavior from standard text-concatenation-based synthesis. Across nine downstream datasets spanning math, science QA, code generation, GAIA, and multi-agent database diagnosis, Parallel-Synthesis matches or outperforms text-based synthesis on seven datasets and remains close on the other two. It also reduces time-to-first-token by 2.5x-11x, suggesting that direct cache-based synthesis is a promising interface for more native and efficient synthesis over parallel agent branches.