论文

多轮商业代理的LLM评估系统设计

Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents

精选理由

论文教你如何构建可靠的多轮代理评估系统,包含具体方法和生产验证结果。

该研究提出了一种集成方法,用于评估多轮商业代理的LLM-as-a-Judge系统。该方法包括评估规范、模块化LLM裁判、保留意图的用户模拟和人工循环治理。生产研究显示,系统级保真度在重复审计中持续提升,人工审查员与自动裁判在共享反馈循环中共同改进,其组合工作流在任务完成设置中表现出最强的描述性能。

原文 · arXiv cs.AI

Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents

Many LLM-as-a-judge evaluations score fixed outputs under a fixed task definition. Production multi-turn business agents instead require a maintained measurement system: correctness depends on business-specific facts and procedures, outcomes emerge across turns, and failures must be attributed to either agent capability or missing business knowledge before they are actionable. We present an integrated methodology spanning evaluation specification, modular LLM judges, intent-preserving user simulation, and human-in-the-loop governance. The specification defines conversation-level end states and actionable failure ownership. Atomic judges share versioned evidence and feed an explicit aggregation graph. The simulator is released only after task-preservation and stability checks. Independent human audits estimate measurement fidelity, renew tiered reference sets, and route disagreements to label correction, guideline revision, or judge improvement. Production studies show that system-level fidelity improved across repeated audits, that human reviewers and automated judges improved together under the shared feedback loop, and that their combined workflow had the strongest descriptive performance in both reported task-completion settings. Because the studies are observational and the human reference itself required revision, these findings demonstrate operational usefulness rather than causal or universal superiority. The contribution is a practical framework for making multi-turn agent measurement reliable, actionable, and maintainable as the evaluated system and its evidence evolve.