论文多源确认精选73°

Nvidia研究:AI模型处理长任务准确率下降

精选理由

Nvidia发现AI处理长任务会变马虎,给每项加编号、分小块处理能提升准确率。

Nvidia测试了7个开源模型在简单重复任务上的表现。128K-token任务的准确率比4K-token任务低62.8%。即使是最佳模型,在最长任务中也只有17.1%的完全正确率。模型在处理没有ID编号的项目时更容易迷失方向。

原文 · rohanpaul_ai

New Nvidia paper: AI models get sloppier as jobs get longer, even inside their context window, so number every item and split big jobs into small chunks.

Model size didn't guarantee reliability on long, repetitive jobs

Picture an agent updating a huge invoice file line by line. It can read the whole file and still skip a line or update the wrong record.

NVIDIA tested 7 open models on simple, repetitive jobs like adding numbers and sorting lists. Average accuracy was 62.8% lower on 128K-token jobs than on 4K-token jobs.

Even the best model got every item right in only 17.1% of the longest jobs. The models seemed to understand the task but lost their place, especially when items had no ID numbers.

If your agent works through long lists, give every item an ID, process them in small batches, and check every line of output.

– arxiv. org/abs/2609.38712

Title: "Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability"