多向量检索策略选错,nDCG@10 从 0.701 暴跌至 0.109

With the same multi-vector model, and the same dataset, nDCG@10 can drop from 0.701 to 0.109 — rough...

精选理由

做向量检索或 RAG 的开发者注意了:多向量检索中策略选择比模型选择更关键,选错策略可能让最好的模型也白费。建议在调优前先测一下 token 向量的分离度,再决定用 TokenANN 还是 LEMUR。

AI 摘要

Milvus 团队在一条推文中揭示了一个关键发现:在多向量检索中,选择错误的近似检索策略比选错模型带来的性能损失更大。他们使用相同的 Jina-ColBERT-v2 模型和 LoTTE 数据集,仅改变第一阶段近似检索策略,结果 TokenANN 策略的 nDCG@10 达到 0.701,而 LEMUR 策略仅为 0.109,差距约 6 倍。原因是不同策略对模型 token 向量的空间分布(分离度)敏感度不同:对于分布分散的模型(如 Jina),TokenANN 和 MUVERA 效果好;对于分布紧凑的模型(如 AnswerAI),LEMUR 更优。研究者可以通过计算 token 向量 MaxSim 得分的标准差来预判策略选择。

原文 · Milvus

With the same multi-vector model, and the same dataset, nDCG@10 can drop from 0.701 to 0.109 — rough...

With the same multi-vector model, and the same dataset, nDCG @10 can drop from 0.701 to 0.109 — roughly a 6x gap. Why? It's because you changed the approximate retrieval strategy. 𝗜𝗻 𝗺𝘂𝗹𝘁𝗶-𝘃𝗲𝗰𝘁𝗼𝗿 𝗿𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹, 𝗽𝗶𝗰𝗸𝗶𝗻𝗴 𝘁𝗵𝗲 𝘄𝗿𝗼𝗻𝗴 𝘀𝘁𝗿𝗮𝘁𝗲𝗴𝘆 𝗰𝗮𝗻 𝗰𝗼𝘀𝘁 𝘆𝗼𝘂 𝗺𝗼𝗿𝗲 𝘁𝗵𝗮𝗻 𝗽𝗶𝗰𝗸𝗶𝗻𝗴 𝘁𝗵𝗲 𝘄𝗿𝗼𝗻𝗴 𝗺𝗼𝗱𝗲𝗹. Multi-vector models like ColBERT turn every token in a document into its own vector. You can't just load millions of token vectors into an ANN index and search, because the score that matters is document-level MaxSim — for each query token, find its closest token in the document, then sum. 𝗦𝗼 𝗲𝘃𝗲𝗿𝘆 𝗮𝗽𝗽𝗿𝗼𝗮𝗰𝗵 𝗿𝘂𝗻𝘀 𝗶𝗻 𝘁𝘄𝗼 𝘀𝘁𝗮𝗴𝗲𝘀: 𝗮𝗻 𝗮𝗽𝗽𝗿𝗼𝘅𝗶𝗺𝗮𝘁𝗲 𝘀𝗲𝗮𝗿𝗰𝗵 𝗽𝗶𝗰𝗸𝘀 𝗰𝗮𝗻𝗱𝗶𝗱𝗮𝘁𝗲 𝗱𝗼𝗰𝘂𝗺𝗲𝗻𝘁𝘀 𝗳𝗶𝗿𝘀𝘁, 𝘁𝗵𝗲𝗻 𝗲𝘅𝗮𝗰𝘁 𝗠𝗮𝘅𝗦𝗶𝗺 𝗿𝗲-𝗿𝗮𝗻𝗸𝘀 𝘁𝗵𝗲𝗺. 𝗧𝗵𝗲 𝘀𝘁𝗿𝗮𝘁𝗲𝗴𝗶𝗲𝘀 𝗼𝗻𝗹𝘆 𝗱𝗶𝗳𝗳𝗲𝗿 𝗶𝗻 𝗵𝗼𝘄 𝘁𝗵𝗲𝘆 𝗱𝗼 𝘁𝗵𝗮𝘁 𝗳𝗶𝗿𝘀𝘁 𝘀𝘁𝗲𝗽. 𝗪𝗲 𝘁𝗲𝘀𝘁𝗲𝗱 𝘁𝗵𝗿𝗲𝗲: 𝗧𝗼𝗸𝗲𝗻𝗔𝗡𝗡 (𝗶𝗻𝗱𝗲𝘅 𝗲𝘃𝗲𝗿𝘆 𝘁𝗼𝗸𝗲𝗻 𝘃𝗲𝗰𝘁𝗼𝗿 𝗱𝗶𝗿𝗲𝗰𝘁𝗹𝘆), 𝗠𝗨𝗩𝗘𝗥𝗔 (𝗿𝗮𝗻𝗱𝗼𝗺-𝗽𝗿𝗼𝗷𝗲𝗰𝘁𝗶𝗼𝗻 𝗰𝗼𝗺𝗽𝗿𝗲𝘀𝘀𝗶𝗼𝗻), 𝗟𝗘𝗠𝗨𝗥 (𝘁𝗿𝗮𝗶𝗻 𝗮𝗻 𝗠𝗟𝗣 𝘁𝗼 𝗰𝗼𝗺𝗽𝗿𝗲𝘀𝘀). On LoTTE, with Jina-ColBERT-v2 held fixed: 📈 TokenANN — nDCG @10 = 0.701 📉 LEMUR — nDCG @10 = 0.109 The only thing that changed is the first-stage approximation. For scale: moving from a plain dense model up to this multi-vector one lifted the exact score from 0.611 to 0.722 on the same data. The loss from the wrong strategy was bigger than the gain from the better model. 𝗪𝗵𝘆? 𝗔 𝘀𝘁𝗿𝗼𝗻𝗴 𝗽𝗿𝗲𝗱𝗶𝗰𝘁𝗼𝗿 𝗶𝘀 𝗲𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴-𝘀𝗽𝗮𝗰𝗲 𝘀𝗲𝗽𝗮𝗿𝗮𝗯𝗶𝗹𝗶𝘁𝘆 — 𝘄𝗵𝗲𝘁𝗵𝗲𝗿 𝗮 𝗺𝗼𝗱𝗲𝗹'𝘀 𝘁𝗼𝗸𝗲𝗻 𝘃𝗲𝗰𝘁𝗼𝗿𝘀 𝘀𝗶𝘁 𝗳𝗮𝗿 𝗮𝗽𝗮𝗿𝘁 𝗼𝗿 𝗽𝗶𝗹𝗲 𝘁𝗼𝗴𝗲𝘁𝗵𝗲𝗿. 𝗙𝗼𝗿 𝘀𝗽𝗿𝗲𝗮𝗱 𝗼𝘂𝘁 (𝗝𝗶𝗻𝗮): each token is a precise probe, so TokenANN and MUVERA land on the right document tokens. LEMUR, though, risks collapsing on long-tailed data. 𝗙𝗼𝗿 𝗰𝗹𝘂𝘀𝘁𝗲𝗿𝗲𝗱 (𝗔𝗻𝘀𝘄𝗲𝗿𝗔𝗜): one query token pulls back a crowd of close-but-irrelevant tokens, so TokenANN and MUVERA produce weak candidates. LEMUR is the one that still works. 𝗬𝗼𝘂 𝗰𝗮𝗻 𝗺𝗲𝗮𝘀𝘂𝗿𝗲 𝘀𝗲𝗽𝗮𝗿𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗯𝗲𝗳𝗼𝗿𝗲 𝗰𝗼𝗺𝗺𝗶𝘁𝘁𝗶𝗻𝗴. Sample a few hundred token vectors, treat each as a query, and take the standard deviation of their MaxSim scores across documents. Jina sits at 0.157; AnswerAI at 0.050 — at 0.050, nearly every document scores the same, so relevant and irrelevant blur. Recall tracks it: Jina's TokenANN holds Math R @100 of 68.5–88.5% across the four datasets, AnswerAI's 44.6–65.5%. 𝗦𝗼 𝗯𝗲𝗳𝗼𝗿𝗲 𝘁𝘂𝗻𝗶𝗻𝗴 𝗮𝗻𝘆𝘁𝗵𝗶𝗻𝗴: → Wide spread (closer to 0.15): start with TokenANN or MUVERA → Tight spread (closer to 0.05): start with LEMUR 𝗧𝗵𝗲 𝗺𝗼𝗱𝗲𝗹 𝗶𝘀 𝗼𝗻𝗹𝘆 𝗵𝗮𝗹𝗳 𝘁𝗵𝗲 𝗱𝗲𝗰𝗶𝘀𝗶𝗼𝗻. 𝗣𝗮𝗶𝗿 𝗶𝘁 𝘄𝗶𝘁𝗵 𝘁𝗵𝗲 𝘄𝗿𝗼𝗻𝗴 𝘀𝘁𝗿𝗮𝘁𝗲𝗴𝘆, 𝗮𝗻𝗱 𝘁𝗵𝗲 𝗯𝗲𝘀𝘁 𝗺𝗼𝗱𝗲𝗹 𝗰𝗮𝗻'𝘁 𝗺𝗮𝗸𝗲 𝘂𝗽 𝘁𝗵𝗲 𝗱𝗶𝗳𝗳𝗲𝗿𝗲𝗻𝗰𝗲. 💬 0 🔄 0 ❤️ 0 👀 27 ⚡