CRAG 打破 RAG 系统错误雪球循环:检索即评估

𝗜𝗻 𝗮 𝗹𝗼𝗻𝗴-𝗿𝘂𝗻𝗻𝗶𝗻𝗴 𝗥𝗔𝗚 𝘀𝘆𝘀𝘁𝗲𝗺, 𝘁𝗵𝗲 𝗺𝗼𝘀𝘁 𝗱𝗮𝗻𝗴𝗲𝗿𝗼𝘂𝘀 𝗯𝘂𝗴 ...

精选理由

做 RAG 系统的开发者最怕错误被反复放大,CRAG 用简单评估机制切断雪球效应,值得在长期运行的生产环境中试试。

AI 摘要

长期运行的 RAG 系统最危险的 bug 不是单次错误答案,而是错误被反复检索、强化,最终被系统当作事实。CRAG(Corrective RAG)通过在检索和生成之间加入轻量级评估步骤,对文档进行置信度评分(0.9 以上直接使用,0.5-0.9 补充网络搜索,低于 0.5 丢弃),并在下次检索前预过滤掉低分内容,从而打破“检索→存储→强化”的恶性循环。CRAG 需要向量数据库支持动态存储置信度、混合检索和分区键,Milvus 原生支持这些能力。

原文 · Milvus

𝗜𝗻 𝗮 𝗹𝗼𝗻𝗴-𝗿𝘂𝗻𝗻𝗶𝗻𝗴 𝗥𝗔𝗚 𝘀𝘆𝘀𝘁𝗲𝗺, 𝘁𝗵𝗲 𝗺𝗼𝘀𝘁 𝗱𝗮𝗻𝗴𝗲𝗿𝗼𝘂𝘀 𝗯𝘂𝗴 ...

𝗜𝗻 𝗮 𝗹𝗼𝗻𝗴-𝗿𝘂𝗻𝗻𝗶𝗻𝗴 𝗥𝗔𝗚 𝘀𝘆𝘀𝘁𝗲𝗺, 𝘁𝗵𝗲 𝗺𝗼𝘀𝘁 𝗱𝗮𝗻𝗴𝗲𝗿𝗼𝘂𝘀 𝗯𝘂𝗴 𝗶𝘀𝗻'𝘁 𝗮 𝘀𝗶𝗻𝗴𝗹𝗲 𝘄𝗿𝗼𝗻𝗴 𝗮𝗻𝘀𝘄𝗲𝗿. 𝗜𝘁'𝘀 𝘁𝗵𝗲 𝗼𝗻𝗲 𝘁𝗵𝗮𝘁 𝘀𝗻𝗼𝘄𝗯𝗮𝗹𝗹𝘀: 𝗿𝗲𝘁𝗿𝗶𝗲𝘃𝗲𝗱 𝗮𝗴𝗮𝗶𝗻 𝗮𝗻𝗱 𝗮𝗴𝗮𝗶𝗻, 𝗿𝗲𝗶𝗻𝗳𝗼𝗿𝗰𝗲𝗱 𝗲𝗮𝗰𝗵 𝘁𝗶𝗺𝗲, 𝘂𝗻𝘁𝗶𝗹 𝘁𝗵𝗲 𝘀𝘆𝘀𝘁𝗲𝗺 𝘁𝗿𝗲𝗮𝘁𝘀 𝗶𝘁 𝗮𝘀 𝗳𝗮𝗰𝘁. It's easy to miss: the model generates from a bad retrieval → no one corrects it, so the system assumes it's right → it gets written back to memory → the next query pulls it up again and reinforces the error. 𝗖𝗥𝗔𝗚 (𝗖𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝘃𝗲 𝗥𝗔𝗚) 𝗯𝗿𝗲𝗮𝗸𝘀 𝘁𝗵𝗶𝘀 𝗹𝗼𝗼𝗽. It adds one step between retrieval and generation: evaluation. A lightweight evaluator judges whether each retrieved document actually answers the question and tags it with a confidence score, stored alongside the memory. The evaluator sorts results into three tiers: • 0.9 → correct: refine and use • 0.5–0.9 → ambiguous: add a web search • <0.5 → wrong: discard and search instead On the next retrieval, the system pre-filters first, keeping only entries above 0.7. Weak content gets screened out before it's reused, so it never re-enters the store → retrieve → reinforce loop. CRAG needs a vector database that can store confidence scores dynamically, run hybrid r milvus.io/blog/fix-rag-r… en #RAG . #VectorDatabase � #Milvus � #LLM � #AIEngineering �𝘁𝗮𝗱𝗮𝘁𝗮 𝗳𝗶𝗹𝘁𝗲𝗿𝗶𝗻𝗴, 𝗵𝘆𝗯𝗿𝗶𝗱 𝗿𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹, 𝗮𝗻𝗱 𝗣𝗮𝗿𝘁𝗶𝘁𝗶𝗼𝗻 𝗞𝗲𝘆 𝗼𝘂𝘁 𝗼𝗳 𝘁𝗵𝗲 𝗯𝗼𝘅, 𝘄𝗵𝗶𝗰𝗵 𝗶𝘀 𝗲𝘅𝗮𝗰𝘁𝗹𝘆 𝘄𝗵𝗮𝘁 𝗖𝗥𝗔𝗚 𝗻𝗲𝗲𝗱𝘀. 𝗟𝗲𝗮𝗿𝗻 𝗵𝗼𝘄 𝘁𝗼 𝘀𝗲𝘁 𝘂𝗽 𝗖𝗥𝗔𝗚: https://t.co/iebI7W1lc6 #RAG #VectorDatabase #Milvus #LLM #AIEngineering 💬 0 🔄 0 ❤️ 1 👀 36 ⚡