这篇论文揭示了AI在金融分析中的检索-整合差距,并探讨了如何通过工作流程设计来解决这个问题。对于对AI在金融领域应用感兴趣的人来说,这是一个值得关注的发现。
研究发现,在长文本金融分析中存在检索-整合差距,即检索到的信息对投资判断的影响降至实验噪声水平。即使是更强大的模型也无法消除这一差距。研究指出,压缩摘要和源文本查找共同将披露信息传递到判断中,而工作流程架构决定了这种传递是否成功。AI分析师的表现由模型能力和工作流程架构共同决定。基于检索的评估可以证明系统在投资判断中忽略了他们实际上检索到的信息。
Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.