论文精选73°

Adobe提出实时数据科学任务智能体评估方法

精选理由

Adobe解决了AI智能体评估中动态数据的问题,用代码基准替代静态描述,评估准确率提升近三成。

Adobe发布论文提出基于技能的智能体评估方法,针对实时数据科学任务。研究团队使用代码而非静态描述作为正确答案基准,在53个测试案例中,AI评估比人工描述准确提高29%,token使用减少16%。当没有正确答案参考时,AI评估表现甚至低于随机水平。

原文 · rohanpaul_ai

Standard agent tests assume the right answer never changes, but on live data it does.

So this Adobe paper checks against code that recalculates it and gets more accurate grades.

Adobe's fix is to write the right answer as code that fetches the current result each time the test runs. An AI grader then checks the agent's reply against it.

On 53 test cases, the AI grader matched human experts 29% better this way than with a written description, and used 16% fewer tokens. With no answer to check against, the AI grader did worse than random.

– arxiv. org/abs/2609.16487

Title: "Skill-based Agentic Evaluation for Real-time Data Science Tasks"