别总怪模型了,很多引文错误出在工程层。这篇文章帮你分清五种引文故障,对症下药。
Milvus团队指出LLM在RAG中频繁引用了不支持的来源。引文失败分为两类:忠实性错误(生成内容与检索文档不符,如模型声称150W功耗但文档只说低功耗)和引文准确性错误(元数据映射错误、缺失引用、幽灵引用、弱支持引用、过度引用)。其中幽灵引用常因索引重建后ID过期导致。修复方案因错误类型而异:忠实性问题调整生成层约束或基座模型,引文准确性问题需工程层修复元数据管理。
𝗟𝗟𝗠𝘀 𝗸𝗲𝗲𝗽 𝗰𝗶𝘁𝗶𝗻𝗴 𝘀𝗼𝘂𝗿𝗰𝗲𝘀 𝘁𝗵𝗮𝘁 𝗱𝗼𝗻'𝘁 𝘀𝗮𝘆 𝘄𝗵𝗮𝘁 𝘁𝗵𝗲𝘆 ...
𝗟𝗟𝗠𝘀 𝗸𝗲𝗲𝗽 𝗰𝗶𝘁𝗶𝗻𝗴 𝘀𝗼𝘂𝗿𝗰𝗲𝘀 𝘁𝗵𝗮𝘁 𝗱𝗼𝗻'𝘁 𝘀𝗮𝘆 𝘄𝗵𝗮𝘁 𝘁𝗵𝗲𝘆 𝗰𝗹𝗮𝗶𝗺. 𝗛𝗲𝗿𝗲'𝘀 𝗵𝗼𝘄 𝘁𝗼 𝗳𝗶𝘅 𝗶𝘁. 𝗧𝗵𝗲 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 𝗶𝘀 𝗳𝗮𝗺𝗶𝗹𝗶𝗮𝗿 𝗳𝗿𝗼𝗺 𝗮𝗻𝘆 𝗵𝗶𝗴𝗵-𝘀𝘁𝗮𝗸𝗲𝘀 𝗥𝗔𝗚: 𝗮 𝗰𝗶𝘁𝗮𝘁𝗶𝗼𝗻 𝘁𝗮𝗴 𝗱𝗼𝗲𝘀𝗻'𝘁 𝗺𝗲𝗮𝗻 𝘁𝗵𝗲 𝗰𝗶𝘁𝗲𝗱 𝗽𝗮𝘀𝘀𝗮𝗴𝗲 𝘀𝘂𝗽𝗽𝗼𝗿𝘁𝘀 𝘁𝗵𝗲 𝗰𝗹𝗮𝗶𝗺. The answer comes back with clean [1][2][3] markers, the right document is in the top-5, everything looks sourced. Then you open the citation and the cited passage says nothing that backs the sentence attached to it. Most teams blame the generation layer and swap the model, rewrite the prompt, pile on instructions. 𝗕𝘂𝘁 𝗮 𝗹𝗼𝘁 𝗼𝗳 𝗰𝗶𝘁𝗮𝘁𝗶𝗼𝗻 𝗲𝗿𝗿𝗼𝗿𝘀 𝗻𝗲𝘃𝗲𝗿 𝗰𝗮𝗺𝗲 𝗳𝗿𝗼𝗺 𝗴𝗲𝗻𝗲𝗿𝗮𝘁𝗶𝗼𝗻. 𝗔 𝗯𝗿𝗼𝗸𝗲𝗻 𝗰𝗶𝘁𝗮𝘁𝗶𝗼𝗻 𝗶𝘀 𝘂𝘀𝘂𝗮𝗹𝗹𝘆 𝗼𝗻𝗲 𝗼𝗳 𝘁𝘄𝗼 𝗽𝗿𝗼𝗯𝗹𝗲𝗺𝘀, 𝗳𝗶𝘅𝗲𝗱 𝗶𝗻 𝗱𝗶𝗳𝗳𝗲𝗿𝗲𝗻𝘁 𝗽𝗹𝗮𝗰𝗲𝘀. • 𝗙𝗮𝗶𝘁𝗵𝗳𝘂𝗹𝗻𝗲𝘀𝘀 is whether the generated text itself is hallucinated. The model says "150W power draw," but the retrieved docs only say "low-power." The 150W is invented. 𝗙𝗶𝘅 𝗶𝘁 𝗶𝗻 𝘁𝗵𝗲 𝗴𝗲𝗻𝗲𝗿𝗮𝘁𝗶𝗼𝗻 𝗹𝗮𝘆𝗲𝗿: tighten prompt constraints, or switch the base model. • 𝗖𝗶𝘁𝗮𝘁𝗶𝗼𝗻 𝗔𝗰𝗰𝘂𝗿𝗮𝗰𝘆 is whether the citation's metadata maps to the right place. The marker points to the wrong passage in the right document, or to a chunk ID that re-chunking already deleted. 𝗙𝗶𝘅 𝗶𝘁 𝗶𝗻 𝘁𝗵𝗲 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝗶𝗻𝗴 𝗹𝗮𝘆𝗲𝗿: citation resolution and metadata management, which has nothing to do with model quality. 𝗖𝗶𝘁𝗮𝘁𝗶𝗼𝗻 𝗔𝗰𝗰𝘂𝗿𝗮𝗰𝘆 𝗳𝗮𝗶𝗹𝘀 𝗶𝗻 𝗳𝗶𝘃𝗲 𝘀𝗵𝗮𝗽𝗲𝘀: • 𝗠𝗶𝘀𝘀𝗶𝗻𝗴 𝗰𝗶𝘁𝗮𝘁𝗶𝗼𝗻: a factual statement with no source attached. Fix the output format and attribution requirements. • 𝗚𝗵𝗼𝘀𝘁 𝗰𝗶𝘁𝗮𝘁𝗶𝗼𝗻: the cited chunk ID no longer exists, usually after an index rebuild leaves old IDs stale while the prompt template keeps emitting them. Fix the ID lifecycle, not the model. • 𝗪𝗲𝗮𝗸-𝘀𝘂𝗽𝗽𝗼𝗿𝘁 𝗰𝗶𝘁𝗮𝘁𝗶𝗼𝗻: the chunk exists but doesn't support the claim. The model says "rate limit is 100 requests/sec" and cites a chunk that only mentions "supports batch calls." Add claim-level entailment checks instead of trusting the marker. • 𝗢𝘃𝗲𝗿-𝗰𝗶𝘁𝗮𝘁𝗶𝗼𝗻: several loosely related sources stacked behind one sentence, which dilutes trust instead of building it. • 𝗗𝗼𝗰-𝗹𝗲𝘃𝗲𝗹 𝗰𝗶𝘁𝗮𝘁𝗶𝗼𝗻: attribution to "Document A" but not the actual passage. For long documents, that's barely verifiable. The first three are the ones to monitor first. Each fails in a different layer, so misclassifying the error sends you to repair the wrong system. Before changing the model, check what kind of citation failure you actually have. 𝗔 𝗥𝗔𝗚 𝘀𝘆𝘀𝘁𝗲𝗺 𝘁𝗵𝗮𝘁 𝘀𝗵𝗼𝘄𝘀 𝗰𝗶𝘁𝗮𝘁𝗶𝗼𝗻 𝗺𝗮𝗿𝗸𝗲𝗿𝘀 𝗯𝘂𝘁 𝘃𝗮𝗹𝗶𝗱𝗮𝘁𝗲𝘀 𝗻𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲𝗺 𝗰𝗮𝗻 𝗯𝗲 𝘄𝗼𝗿𝘀𝗲 𝘁𝗵𝗮𝗻 𝗼𝗻𝗲 𝘄𝗶𝘁𝗵 𝗻𝗼 𝗰𝗶𝘁𝗮𝘁𝗶𝗼𝗻𝘀 𝗮𝘁 𝗮𝗹𝗹. 𝗔 𝗰𝗹𝗲𝗮𝗻 [𝟭][𝟮][𝟯] 𝘀𝘂𝗿𝗳𝗮𝗰𝗲 𝗺𝗮𝗸𝗲𝘀 𝘀𝘁𝗮𝗹𝗲, 𝘄𝗿𝗼𝗻𝗴, 𝗼𝗿 𝘂𝗻𝘀𝘂𝗽𝗽𝗼𝗿𝘁𝗲𝗱 𝗲𝘃𝗶𝗱𝗲𝗻𝗰𝗲 𝗹𝗼𝗼𝗸 𝘁𝗿𝘂𝘀𝘁𝘄𝗼𝗿𝘁𝗵𝘆. 💬 0 🔄 0 ❤️ 0 👀 66 ⚡