技巧精选

RAG评估陷阱:单一平均分可能掩盖幻觉,试试声明级评估

𝗔 𝘀𝗶𝗻𝗴𝗹𝗲 𝟭–𝟱 𝘀𝗰𝗼𝗿𝗲 𝗶𝘀 𝗮 𝗯𝗮𝗱 𝘄𝗮𝘆 𝘁𝗼 𝗷𝘂𝗱𝗴𝗲 𝗥𝗔𝗚 𝗾𝘂𝗮𝗹𝗶𝘁𝘆 𝗶𝗻 ...

精选理由

如果你在用RAG做生产系统,这篇讲透了为什么平均分不靠谱,还给了按声明颗粒度和问题类型精准监测的方法,连Milvus怎么分桶都说了,很实用。

AI 摘要

单个1-5分的RAG质量评分会隐藏严重问题:一个回答90%基于文档,但10%虚构核心参数就不可用,平均分仍显示4分。幻觉分布也不均匀,数值查找或多条件问题类型的幻觉率远高于平均,不按类型分桶就看不到偏差。优化答案相关性时,添加提示词“提供更完整背景”可能提升相关度但导致模型依赖参数知识,降低忠实度。更可靠的方法是声明级评估:将回答拆成原子事实,用NLI模型检查每个声明是否被检索内容支撑,计算接地率,并对关键参数设置硬性阻断。按问题类型分桶评分,Milvus可用标量字段直接过滤分析,不依赖额外报表管线。

原文 · Milvus

𝗔 𝘀𝗶𝗻𝗴𝗹𝗲 𝟭–𝟱 𝘀𝗰𝗼𝗿𝗲 𝗶𝘀 𝗮 𝗯𝗮𝗱 𝘄𝗮𝘆 𝘁𝗼 𝗷𝘂𝗱𝗴𝗲 𝗥𝗔𝗚 𝗾𝘂𝗮𝗹𝗶𝘁𝘆 𝗶𝗻 ...

𝗔 𝘀𝗶𝗻𝗴𝗹𝗲 𝟭–𝟱 𝘀𝗰𝗼𝗿𝗲 𝗶𝘀 𝗮 𝗯𝗮𝗱 𝘄𝗮𝘆 𝘁𝗼 𝗷𝘂𝗱𝗴𝗲 𝗥𝗔𝗚 𝗾𝘂𝗮𝗹𝗶𝘁𝘆 𝗶𝗻 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻. 𝗛𝗲𝗿𝗲'𝘀 𝘄𝗵𝘆. 𝗙𝗶𝗿𝘀𝘁, 𝗮𝗻 𝗼𝘃𝗲𝗿𝗮𝗹𝗹 𝗮𝘃𝗲𝗿𝗮𝗴𝗲 𝗵𝗶𝗱𝗲𝘀 𝗵𝗶𝗴𝗵-𝗿𝗶𝘀𝗸 𝗵𝗮𝗹𝗹𝘂𝗰𝗶𝗻𝗮𝘁𝗶𝗼𝗻𝘀. A long answer can be 90% grounded, but if the other 10% fabricates a core parameter, it's unusable in finance, medicine, or law. The averaged score still reads 4 out of 5, and the team assumes the system is basically fine. 𝗦𝗲𝗰𝗼𝗻𝗱, 𝗵𝗮𝗹𝗹𝘂𝗰𝗶𝗻𝗮𝘁𝗶𝗼𝗻𝘀 𝗮𝗿𝗲𝗻'𝘁 𝗲𝘃𝗲𝗻𝗹𝘆 𝗱𝗶𝘀𝘁𝗿𝗶𝗯𝘂𝘁𝗲𝗱. Average Groundedness can look healthy while certain query types, like numeric lookups or multi-condition questions, hallucinate far above the mean. Without bucketing by question type, that skew stays invisible. 𝗧𝗵𝗶𝗿𝗱, 𝗺𝗲𝘁𝗿𝗶𝗰𝘀 𝗰𝗮𝗻 𝗺𝗮𝘀𝗸 𝗲𝗮𝗰𝗵 𝗼𝘁𝗵𝗲𝗿 When you optimize for Answer Relevance and add prompt instructions like "give fuller background," relevance scores do go up, but the model starts pulling on parametric knowledge to fill evidence gaps, and Faithfulness can drop without showing up in the headline score. One metric rises while another degrades out of sight. That's why Groundedness isn't a one-time check: it has to be monitored after every prompt change. 𝗧𝗵𝗲 𝗺𝗼𝗿𝗲 𝗿𝗲𝗹𝗶𝗮𝗯𝗹𝗲 𝗮𝗽𝗽𝗿𝗼𝗮𝗰𝗵 𝗶𝘀 𝗰𝗹𝗮𝗶𝗺-𝗹𝗲𝘃𝗲𝗹 𝗲𝘃𝗮𝗹𝘂𝗮𝘁𝗶𝗼𝗻, 𝗶𝗻 𝗳𝗼𝘂𝗿 𝘀𝘁𝗲𝗽𝘀: • 𝗔𝘁𝗼𝗺𝗶𝗰 𝗰𝗹𝗮𝗶𝗺 𝗲𝘅𝘁𝗿𝗮𝗰𝘁𝗶𝗼𝗻: break the full answer into independent factual assertions. "The XR-2048 is a low-power device, draws 150W, and supports hot-swapping" becomes three separate claims. • 𝗘𝗻𝘁𝗮𝗶𝗹𝗺𝗲𝗻𝘁 𝗰𝗵𝗲𝗰𝗸: for each claim, test whether the retrieved context actually supports it. Use an NLI model or an LLM-as-judge. • 𝗤𝘂𝗮𝗻𝘁𝗶𝗳𝘆: Groundedness = supported claims / total claims. • 𝗖𝗿𝗶𝘁𝗶𝗰𝗮𝗹-𝗲𝗿𝗿𝗼𝗿 𝗴𝗮𝘁𝗶𝗻𝗴: in high-stakes work, if a claim covering a key parameter is unsupported, block the answer regardless of how high the overall score is. 𝗕𝘂𝗰𝗸𝗲𝘁 𝘆𝗼𝘂𝗿 𝘀𝗰𝗼𝗿𝗲𝘀 𝗯𝘆 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻 𝘁𝘆𝗽𝗲 𝗶𝗻𝘀𝘁𝗲𝗮𝗱 𝗼𝗳 𝗿𝗲𝗮𝗱𝗶𝗻𝗴 𝗼𝗻𝗲 𝗴𝗹𝗼𝗯𝗮𝗹 𝗮𝘃𝗲𝗿𝗮𝗴𝗲. A per-claim view tells you more than whether the answer looks solid overall. It points to the exact sentence that has nothing behind it, and it makes debugging far easier, because you can trace every unsupported claim back to where it broke: retrieval never pulled the evidence, the reranker buried it, the context window cut it off, or the model just ignored what it had. 𝗕𝘂𝗰𝗸𝗲𝘁𝗶𝗻𝗴 𝗶𝘀 𝘀𝘁𝗿𝗮𝗶𝗴𝗵𝘁𝗳𝗼𝗿𝘄𝗮𝗿𝗱 𝘁𝗼 𝘄𝗶𝗿𝗲 𝘂𝗽. Tag each evaluation record with its query type, then filter on that tag. KStore evaluation records with fields like query_type, claim_count, and unsupported_claim_count. 𝗠𝗶𝗹𝘃𝘂𝘀 can do this with a scalar or dynamic field, so the per-bucket numbers fall out of a single filter. That lets you slice results without building a separate reporting pipeline. 𝗜𝗳 𝘆𝗼𝘂𝗿 𝗥𝗔𝗚 𝘀𝘆𝘀𝘁𝗲𝗺 𝗿𝗲𝗽𝗼𝗿𝘁𝘀 𝗮 𝘀𝗶𝗻𝗴𝗹𝗲 𝗾𝘂𝗮𝗹𝗶𝘁𝘆 𝘀𝗰𝗼𝗿𝗲, 𝗰𝗵𝗲𝗰𝗸 𝘄𝗵𝗮𝘁 𝗮 𝗽𝗲𝗿-𝗰𝗹𝗮𝗶𝗺, 𝗽𝗲𝗿-𝗯𝘂𝗰𝗸𝗲𝘁 𝗯𝗿𝗲𝗮𝗸𝗱𝗼𝘄𝗻 𝗱𝗼𝗲𝘀 𝘁𝗼 𝗶𝘁. 𝗧𝗵𝗲 𝗻𝘂𝗺𝗯𝗲𝗿 𝘁𝗵𝗮𝘁 𝗹𝗼𝗼𝗸𝗲𝗱 𝗳𝗶𝗻𝗲 𝘂𝘀𝘂𝗮𝗹𝗹𝘆 𝗱𝗼𝗲𝘀𝗻'𝘁 𝘀𝘂𝗿𝘃𝗶𝘃𝗲. 💬 0 🔄 0 ❤️ 0 👀 53 ⚡