做 RAG 的团队别再盲目换大模型了——Milvus 这篇诊断法帮你精准定位检索瓶颈,从精确术语到长尾查询都有对应解法,建议直接收藏实操。
当 RAG 系统给出错误答案时,团队通常第一时间换更大的模型或调 prompt,但 Milvus 团队指出,真正该先修的是检索环节。他们提出一个三步诊断法:先按查询类型(精确术语、多跳、长尾、不可回答)构建黄金测试集,然后按桶计算 Recall@k,最后根据弱桶定位问题——精确术语桶低说明稠密检索对精确字符串有盲点,应加混合搜索;多跳桶低说明答案被切分或候选集太小;长尾桶低说明用户措辞与文档术语不匹配,需加查询改写;所有桶都低则说明嵌入模型不适合领域。这种方法能精准定位检索失败的具体原因,而非笼统地认为“召回率差”。
When RAG gives a wrong answer, the team's first move is usually a bigger model or more prompt tweaki...
When RAG gives a wrong answer, the team's first move is usually a bigger model or more prompt tweaking. 𝗕𝘂𝘁 𝘁𝗵𝗲 𝗳𝗶𝗿𝘀𝘁 𝘁𝗵𝗶𝗻𝗴 𝘁𝗼 𝗳𝗶𝘅 𝗶𝘀 𝘂𝘀𝘂𝗮𝗹𝗹𝘆 𝗿𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹. 𝗛𝗲𝗿𝗲'𝘀 𝗮 𝘁𝗵𝗿𝗲𝗲-𝘀𝘁𝗲𝗽 𝘄𝗮𝘆 𝘁𝗼 𝗺𝗲𝗮𝘀𝘂𝗿𝗲 𝘄𝗵𝗲𝗿𝗲 𝗶𝘁'𝘀 𝗮𝗰𝘁𝘂𝗮𝗹𝗹𝘆 𝗯𝗿𝗲𝗮𝗸𝗶𝗻𝗴: 1️⃣ 𝗕𝘂𝗶𝗹𝗱 𝗮 𝘀𝗲𝘁 𝗼𝗳 𝗴𝗼𝗹𝗱𝗲𝗻 𝗾𝘂𝗲𝗿𝗶𝗲𝘀, 𝗯𝘂𝗰𝗸𝗲𝘁𝗲𝗱 𝗯𝘆 𝗾𝘂𝗲𝗿𝘆 𝘁𝘆𝗽𝗲. 5–10 per bucket: exact terms (model numbers, API names, error codes), multi-hop, long-tail, and unanswerable. 2️⃣ 𝗥𝘂𝗻 𝘁𝗵𝗲𝗺 𝗮𝗻𝗱 𝗰𝗵𝗲𝗰𝗸 𝗥𝗲𝗰𝗮𝗹𝗹@𝗸 𝗽𝗲𝗿 𝗯𝘂𝗰𝗸𝗲𝘁. Don't look at the overall average. It hides the weak bucket. 3️⃣ 𝗧𝗵𝗲 𝘄𝗲𝗮𝗸 𝗯𝘂𝗰𝗸𝗲𝘁 𝘂𝘀𝘂𝗮𝗹𝗹𝘆 𝘁𝗲𝗹𝗹𝘀 𝘆𝗼𝘂 𝘄𝗵𝗲𝗿𝗲 𝘁𝗼 𝗹𝗼𝗼𝗸 𝗳𝗶𝗿𝘀𝘁: 📦 𝗘𝘅𝗮𝗰𝘁-𝘁𝗲𝗿𝗺 𝗯𝘂𝗰𝗸𝗲𝘁 𝗹𝗼𝘄 → dense retrieval's blind spot for exact strings. Model numbers and error codes often do not land near the right context in embedding space → add hybrid search (dense + BM25). 🔗 𝗠𝘂𝗹𝘁𝗶-𝗵𝗼𝗽 𝗹𝗼𝘄 → the answer got split across chunks, or your candidate set is too small → check your chunking and how many results you pull back 🌊 𝗟𝗼𝗻𝗴-𝘁𝗮𝗶𝗹 𝗹𝗼𝘄 → users' wording doesn't match the docs' terminology → add query rewriting, or add the terminology users actually search for to the docs. 📉 𝗘𝘃𝗲𝗿𝘆𝘁𝗵𝗶𝗻𝗴 𝗹𝗼𝘄 → your embedding model doesn't fit your domain → swap it or fine-tune. 𝗧𝗵𝗮𝘁'𝘀 𝘁𝗵𝗲 𝘃𝗮𝗹𝘂𝗲 𝗼𝗳 𝗱𝗶𝗮𝗴𝗻𝗼𝘀𝗶𝗻𝗴 𝗯𝘆 𝗯𝘂𝗰𝗸𝗲𝘁: 𝗻𝗼𝘁 "𝗿𝗲𝗰𝗮𝗹𝗹 𝗶𝘀 𝗯𝗮𝗱," 𝗯𝘂𝘁 𝘄𝗵𝗶𝗰𝗵 𝗾𝘂𝗲𝗿𝗶𝗲𝘀 𝗳𝗮𝗶𝗹, 𝘄𝗵𝘆, 𝗮𝗻𝗱 𝘄𝗵𝗮𝘁 𝘁𝗼 𝗳𝗶𝘅. 💬 0 🔄 0 ❤️ 0 👀 24 ⚡