论文多源确认精选76°

NVIDIA研究长程智能体错误率问题

精选理由

NVIDIA这篇论文揭示了长程智能体在处理任务时准确率下降的具体数据,对构建可靠AI助手很有参考价值。

NVIDIA发布论文研究长程智能体在128K上下文下的错误率问题。研究显示,当上下文从4K增加到128K时,七个开源模型的准确率下降62.8%。当输入格式变化时,准确率下降36.5%;当每步操作难度增加时,准确率下降39.9%。该研究测量了智能体在处理长表格或账本时失去位置的原因。

原文 · DAIR.AI

Banger paper from NVIDIA on long running agents.

A model can accept 128K tokens of context and still make more mistakes the longer it works through a task.

If your agent loses its place partway through a long table or ledger, this work measures what causes it.

The setup:

Long-Transduction asks a model to keep reading, updating and outputting state-dependent results over thousands of outputs, and varies three factors separately.

Results:

Across seven open-weight models, accuracy drops 62.8% when context grows from 4K to 128K, 36.5% when only the input format changes, and 39.9% when the per-step operation gets harder.

Paper: https://t.co/kjWSHzoegq