做向量检索的团队常遇到多向量模型部署后效果反而不如稠密检索的困惑,Milvus 这篇分析直接点出了根本原因和适用场景,建议做搜索和 RAG 的开发者仔细看看,能帮你避免选型踩坑。
Milvus 团队发文解释了多向量模型在基准测试中表现优异,但在生产环境中效果不如稠密检索的原因。核心问题在于多向量模型使用精确的 MaxSim 评分(每个查询 token 与文档所有 token 比较),而生产环境只能使用近似搜索。稠密检索的近似算法(如 HNSW、IVF)成熟度高,能紧密跟踪精确结果;多向量模型的近似搜索则因压缩或聚合表示导致候选集遗漏,损失更大。实验表明,短文档和简单查询下稠密检索更优,长文档和复杂查询下多向量才值得使用。
In this previous post (https://t.co/77L5mrn5q7), we talked about why multi-vector models sometimes h...
In this previous post ( x.com/milvusio/statu… ), we talked about why multi-vector models sometimes have worse results than plain dense retrieval. 𝗠𝘂𝗹𝘁𝗶-𝘃𝗲𝗰𝘁𝗼𝗿 𝗿𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹 𝘄𝗶𝗻𝘀 𝗼𝗻 𝗯𝗲𝗻𝗰𝗵𝗺𝗮𝗿𝗸𝘀. 𝗜𝗻 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻, 𝗶𝘁 𝗼𝗳𝘁𝗲𝗻 𝗹𝗼𝘀𝗲𝘀 𝘁𝗼 𝗽𝗹𝗮𝗶𝗻 𝗱𝗲𝗻𝘀𝗲 𝗿𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹. 𝗛𝗲𝗿𝗲'𝘀 𝘄𝗵𝘆, 𝗮𝗻𝗱 𝘄𝗵𝗲𝗻 𝘁𝗵𝗲 𝗮𝗱𝘃𝗮𝗻𝘁𝗮𝗴𝗲 𝗮𝗰𝘁𝘂𝗮𝗹𝗹𝘆 𝘀𝘂𝗿𝘃𝗶𝘃𝗲𝘀. 𝗧𝗵𝗲 𝗮𝗻𝘀𝘄𝗲𝗿 𝘀𝘁𝗮𝗿𝘁𝘀 𝘄𝗶𝘁𝗵 𝘄𝗵𝗮𝘁'𝘀 𝗮𝗰𝘁𝘂𝗮𝗹𝗹𝘆 𝗯𝗲𝗶𝗻𝗴 𝗺𝗲𝗮𝘀𝘂𝗿𝗲𝗱: brute-force exact scoring — every query token compared against every document token. Under those conditions, multi-vector almost always beats dense. The gap is modest on simple text datasets, larger on complex ones, and very large on multimodal data. Production can't do exact scoring; corpora are too large. Both dense and multi-vector fall back to approximate search — but the approximation gap isn't symmetric. Dense ANN (HNSW, IVF) is mature enough to closely track brute-force. Multi-vector approximate search isn't: it uses compressed or aggregated representations to generate candidates first, then re-scores only those. The approximation loss is much larger, and documents that would have ranked high under exact MaxSim often don't make the candidate shortlist at all. 𝗧𝗵𝗲 𝗱𝗮𝘁𝗮 𝘀𝗵𝗼𝘄𝘀 𝗲𝘅𝗮𝗰𝘁𝗹𝘆 𝘄𝗵𝗲𝗿𝗲 𝘁𝗵𝗲𝗿𝗲 𝗶𝘀 𝗼𝗿 𝗶𝘀𝗻'𝘁 𝗹𝗼𝘀𝘀. 𝗛𝗼𝘄 𝗺𝘂𝗰𝗵 𝗼𝗳 𝘁𝗵𝗲 𝗲𝘅𝗮𝗰𝘁-𝘀𝗰𝗼𝗿𝗶𝗻𝗴 𝗮𝗱𝘃𝗮𝗻𝘁𝗮𝗴𝗲 𝘀𝘂𝗿𝘃𝗶𝘃𝗲𝘀 𝗱𝗲𝗽𝗲𝗻𝗱𝘀 𝗼𝗻 𝗱𝗼𝗰𝘂𝗺𝗲𝗻𝘁 𝗹𝗲𝗻𝗴𝘁𝗵 𝗮𝗻𝗱 𝗾𝘂𝗲𝗿𝘆 𝗰𝗼𝗺𝗽𝗹𝗲𝘅𝗶𝘁𝘆: 📄 MS MARCO (avg. 87 tokens/doc) — ~1pt lead under exact scoring. Approximate search erased it entirely. 🔬 SciFact (avg. 360 tokens/doc) — 3–9pt exact lead. Approximate search erased it too. Uniformly long documents give token-level matching less to exploit than you'd expect. 🧬 TREC-COVID (avg. 236 tokens/doc, complex biomedical queries) — 12–17pt exact lead. 7–12pts survived. Long docs plus hard queries gave MaxSim enough signal to hold up. 📊 LoTTE (heavy length tail, up to 4,000 tokens/doc) — results split sharply by model and strategy. Length alone doesn't explain it. 𝗦𝗼 𝘄𝗵𝗲𝗻 𝘁𝗼 𝘂𝘀𝗲 𝘄𝗵𝗶𝗰𝗵? 𝗧𝗵𝗲 𝗿𝘂𝗹𝗲 𝗼𝗳 𝘁𝗵𝘂𝗺𝗯: → Short docs, simple queries: dense retrieval wins. Skip the complexity. → Long docs, complex queries: multi-vector holds up. The gain is worth it. → Mixed-length corpus: your model and strategy pairing determines the outcome — more on that next. Milvus @milvusio Sometimes, when teams deploy a multi-vector model, their results come back worse than plain dense retrieval. 𝗧𝗵𝗲 𝗰𝘂𝗹𝗽𝗿𝗶𝘁 𝗶𝘀 𝗮𝗹𝗺𝗼𝘀𝘁 𝗮𝗹𝘄𝗮𝘆𝘀 𝗮 𝗺𝗶𝘀𝗺𝗮𝘁𝗰𝗵 𝗯𝗲𝘁𝘄𝗲𝗲𝗻 𝗵𝗼𝘄 𝗺𝘂𝗹𝘁𝗶-𝘃𝗲𝗰𝘁𝗼𝗿 𝗺𝗼𝗱𝗲𝗹𝘀 𝘀𝗰𝗼𝗿𝗲 𝗱𝗼𝗰𝘂𝗺𝗲𝗻𝘁𝘀 𝗮𝗻𝗱 𝘄𝗵𝗮𝘁 𝘁𝗵𝗲𝗶𝗿 𝗶𝗻𝗱𝗲𝘅 𝗰𝗮𝗻 𝗮𝗰𝘁𝘂𝗮𝗹𝗹𝘆 𝗰𝗼𝗺𝗽𝘂𝘁𝗲. Multi-vector models produce one vector per token. A 200-token document becomes 200 index entries. Standard indexes can't handle this because they're built to find the nearest vector, not the nearest document. Finding the nearest document requires a different operation: for each query token, find its best match among all tokens in that document, then sum those matches. This process is called MaxSim. A standard index lookup can't do that — it has no concept of which vectors belong together. MaxSim needs to aggregate across an entire document's worth of vectors simultaneously. The index may know which document each vector came from, but its native scoring is still vector-level, not document-level MaxSim. 𝗦𝗼 𝘆𝗼𝘂 𝗻𝗲𝗲𝗱 𝗮 𝗯𝗿𝗶𝗱𝗴𝗲 𝗯𝗲𝘁𝘄𝗲𝗲𝗻 𝗵𝗼𝘄 𝗺𝘂𝗹𝘁𝗶-𝘃𝗲𝗰𝘁𝗼𝗿 𝗺𝗼𝗱𝗲𝗹𝘀 𝘀𝗰𝗼𝗿𝗲 𝗿𝗲𝗹𝗲𝘃𝗮𝗻𝗰𝗲 𝗮𝗻𝗱 𝘄𝗵𝗮𝘁 𝘀𝘁𝗮𝗻𝗱𝗮𝗿𝗱 𝗶𝗻𝗱𝗲𝘅𝗲𝘀 𝗰𝗮𝗻 𝗮𝗰𝘁𝘂𝗮𝗹𝗹𝘆 𝗰𝗼𝗺𝗽𝘂𝘁𝗲. 𝗧𝗵𝗿𝗲𝗲 𝘀𝘁𝗿𝗮𝘁𝗲𝗴𝗶𝗲𝘀 𝗲𝘅𝗶𝘀𝘁 𝘁𝗼 𝗯𝗿𝗶𝗱𝗴𝗲 𝘁𝗵𝗲 𝗶𝘀𝘀𝘂𝗲. They're mutually exclusive, and each then feeds into the same second stage: exact reranking over the candidate pool using MaxSim. The difference is how each strategy compresses the multi-vector representation to make that first stage fast. 🗂 𝗧𝗼𝗸𝗲𝗻𝗔𝗡𝗡 • No compression — every token vector goes directly into the index • One ANN search per query token, results aggregated back to documents • Index grows 200× for a 200-token document Works best when your model produces highly discriminative token vectors — meaning individual tokens are distinctive enough that ANN can reliably find the right ones. High fidelity. Highest latency and storage cost of the three. 📐 𝗠𝗨𝗩𝗘𝗥𝗔 • Compresses all token vectors into one fixed-length vector per document via random projections • No training required, no corpus-specific tuning • Plugs into standard ANN search like any single-vector system The catch: those fixed-length vectors can exceed 14,000 dimensions at typical configurations. High-dimensional ANN has its own efficiency problem — product quantization is the standard way to compress them back down. Consistent recall across datasets with minimal tuning makes this the most portable option. 🧠 𝗟𝗘𝗠𝗨𝗥 • Trains a small network to produce one compressed vector per document • Training signal is MaxSim scores — it learns what "relevant" looks like for your specific corpus • Tends to outperform the other two when your model's token vectors are less distinctive Requires retraining or refitting when the corpus or embedding distribution changes substantially. Carries a structural flaw that tuning won't fix: on datasets where document length varies widely, the training signal conflates length with relevance. Longer documents score higher not because they're more relevant but because a larger token pool gives the MaxSim function more chances to hit a high match. On such datasets, rankings collapse. The candidate pool ratio — how many documents you retrieve before reranking — is the primary dial for trading quality against latency, and it dominates the other architectural choices. Optimize that before you optimize anything else. 𝗧𝗵𝗲 𝗿𝗶𝗴𝗵𝘁 𝘀𝘁𝗿𝗮𝘁𝗲𝗴𝘆 𝗶𝘀𝗻'𝘁 𝗱𝗲𝘁𝗲𝗿𝗺𝗶𝗻𝗲𝗱 𝗯𝘆 𝗶𝗻𝗱𝗲𝘅 𝘀𝗶𝘇𝗲. 𝗜𝘁'𝘀 𝗱𝗲𝘁𝗲𝗿𝗺𝗶𝗻𝗲𝗱 𝗯𝘆 𝘆𝗼𝘂𝗿 𝗺𝗼𝗱𝗲𝗹'𝘀 𝗲𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴 𝘀𝗽𝗮𝗰𝗲, 𝘆𝗼𝘂𝗿 𝗹𝗮𝘁𝗲𝗻𝗰𝘆 𝗯𝘂𝗱𝗴𝗲𝘁, 𝗮𝗻𝗱 𝘄𝗵𝗲𝘁𝗵𝗲𝗿 𝘆𝗼𝘂𝗿 𝗱𝗼𝗰𝘂𝗺𝗲𝗻𝘁𝘀 𝗮𝗿𝗲 𝗹𝗼𝗻𝗴 𝗲𝗻𝗼𝘂𝗴𝗵 𝗳𝗼𝗿 𝗟𝗘𝗠𝗨𝗥'𝘀 𝗹𝗲𝗻𝗴𝘁𝗵 𝗯𝗶𝗮𝘀 𝘁𝗼 𝗯𝗲𝗰𝗼𝗺𝗲 𝗮 𝗹𝗶𝗮𝗯𝗶𝗹𝗶𝘁𝘆. 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 3 👀 328 📊 1 ⚡