PARSER:并行读取深度推理的长上下文智能体
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
PARSER让长文档处理不再受顺序限制,证据位置变化不影响结果,推理速度提升11倍。
PARSER是一种新型长上下文智能体架构,通过并行读取文档和深度推理解耦阅读与推理过程。该模型使用4B参数 backbone 在7K到896K tokens的多跳问答任务中,平均比最强顺序记忆基线高5.7分,在896K tokens时领先12.0分。扩展到9B参数后,PARSER超越DeepSeek-V4-Pro 6.3分,推理延迟最高减少11倍。
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency linearly to document length. We introduce PARSER, which decouples reading from reasoning. A bank of lightweight subagents each bound to a single chunk read the entire document in parallel, while a lead agent reasons in depth through iterative scatter--gather rounds: at each round it broadcasts a query to all subagents, aggregates the returned evidence, and formulates a deeper follow-up query conditioned on what has been found so far. This decoupled design concentrates all learnable behavior in the lead agent, which is optimized with reinforcement learning, while the subagents remain frozen off-the-shelf models. On multi-hop QA with contexts ranging from 7K to 896K tokens, PARSER with a 4B backbone outperforms the strongest sequential memory baseline by 5.7 points on average and by 12.0 points at 896K tokens. Scaling to a 9B backbone, PARSER surpasses DeepSeek-V4-Pro by 6.3 points. Controlled experiments confirm that PARSER is robust to perturbations in evidence position, order, and distance, conditions that cause large accuracy swings in sequential methods, while reducing inference latency by up to 11x.