Milvus用ColQwen和Qwen3-VL-Embedding做了对比,发现多向量在检索带图表的文档时比稠密向量强近18个点,近似搜索不掉分。处理PDF或扫描件可以关注这个结果。
Milvus在DocVQA上对比ColQwen(多向量)与Qwen3-VL-Embedding(稠密)的检索性能。精确搜索下,ColQwen3的nDCG@10为0.698,比稠密的0.521高17.7个百分点。近似搜索(LEMUR,ratio=5.0)中,ColQwen3得0.704,领先18.3点,且近似损失几乎为零。在MS MARCO等文本基准上,多向量优势被近似搜索抹平。多向量通过保留表格、图表等空间结构获得提升,适合发票、报告等视觉文档。
In text retrieval, multi-vector's advantage is conditional. On short documents and simple queries, a...
In text retrieval, multi-vector's advantage is conditional. On short documents and simple queries, approximation usually erases its lead over dense. Visual documents flip that: multi-vector beat the dense baseline by 𝗻𝗲𝗮𝗿𝗹𝘆 𝟭𝟴 𝗽𝗼𝗶𝗻𝘁𝘀, and the lead held almost intact through approximation. We ran ColQwen (multi-vector) against Qwen3-VL-Embedding (dense) on DocVQA. 𝗧𝗵𝗶𝘀 𝘁𝗮𝘀𝗸 𝗶𝘀 𝗺𝘂𝗹𝘁𝗶𝗺𝗼𝗱𝗮𝗹: 𝘁𝗵𝗲 𝗾𝘂𝗲𝗿𝘆 𝗶𝘀 𝘁𝗲𝘅𝘁, 𝘄𝗵𝗶𝗹𝗲 𝘁𝗵𝗲 𝗱𝗼𝗰𝘂𝗺𝗲𝗻𝘁 𝗶𝘀 𝗮𝗻 𝗶𝗺𝗮𝗴𝗲, 𝗮𝗻𝗱 𝘁𝗵𝗲 𝗺𝗼𝗱𝗲𝗹 𝗲𝗻𝗰𝗼𝗱𝗲𝘀 𝘁𝗵𝗲 𝗶𝗺𝗮𝗴𝗲 𝗶𝗻𝘁𝗼 𝘃𝗲𝗰𝘁𝗼𝗿𝘀. For ColBERT-style models, such as ColQwen, the image is encoded into multiple vectors at the patch level. For dense models, such as Qwen3-VL-Embedding, the image is encoded into a single vector. 𝗘𝘅𝗮𝗰𝘁 𝘀𝗲𝗮𝗿𝗰𝗵 (BruteForce, the quality ceiling): • Dense: 𝟬.𝟱𝟮𝟭 nDCG @10 • ColQwen2: 𝟬.𝟲𝟭𝟬 (+8.9 pt) • ColQwen3: 𝟬.𝟲𝟵𝟴 (+17.7 pt) 𝗔𝗽𝗽𝗿𝗼𝘅𝗶𝗺𝗮𝘁𝗲 𝘀𝗲𝗮𝗿𝗰𝗵 (LEMUR, ratio=5.0): • ColQwen2: 𝟬.𝟱𝟵𝟲 (+7.5 pt over dense) • ColQwen3: 𝟬.𝟳𝟬𝟰 (+18.3 pt over dense) 𝗔𝗽𝗽𝗿𝗼𝘅𝗶𝗺𝗮𝘁𝗶𝗼𝗻 𝗰𝗼𝘀𝘁 𝘄𝗮𝘀 𝗲𝗳𝗳𝗲𝗰𝘁𝗶𝘃𝗲𝗹𝘆 𝘇𝗲𝗿𝗼. ColQwen3 landed 0.006 above its own exact score, which is noise on 100 queries. On text benchmarks like MS MARCO and SciFact, approximation usually wipes the multi-vector lead out. Here it didn't move. 𝗧𝗵𝗲 𝗴𝗮𝗽 𝗰𝗼𝗺𝗲𝘀 𝗳𝗿𝗼𝗺 𝘁𝗵𝗲 𝗽𝗮𝗴𝗲 𝗶𝘁𝘀𝗲𝗹𝗳: 𝘁𝗮𝗯𝗹𝗲𝘀, 𝗰𝗵𝗮𝗿𝘁𝘀, 𝗺𝘂𝗹𝘁𝗶-𝗰𝗼𝗹𝘂𝗺𝗻 𝗹𝗮𝘆𝗼𝘂𝘁𝘀, 𝗳𝗼𝗻𝘁 𝗵𝗶𝗲𝗿𝗮𝗿𝗰𝗵𝘆. A dense model folds all of that into one vector and loses the spatial structure. Patch-level multi-vectors keep each region as its own vector, so MaxSim can match a query against the specific parts of the page that matter. A single vector can't carry the structure of a full page. 𝗢𝗻𝗲 𝗰𝗮𝘃𝗲𝗮𝘁: 𝗧𝗼𝗸𝗲𝗻𝗔𝗡𝗡 𝗱𝗼𝗲𝘀𝗻'𝘁 𝘄𝗼𝗿𝗸 𝗵𝗲𝗿𝗲 𝗮𝘁 𝗮𝗹𝗹. Both models scored under 0.05. Visual models emit thousands of patches per page (ColQwen2: 5,143; ColQwen3: 1,250), and each patch covers so little of the image that per-patch ANN can't locate the right page. The usable strategies are the compression ones, MUVERA and LEMUR. 𝗜𝗳 𝘆𝗼𝘂'𝗿𝗲 𝗿𝗲𝘁𝗿𝗶𝗲𝘃𝗶𝗻𝗴 𝗶𝗻𝘃𝗼𝗶𝗰𝗲𝘀, 𝗿𝗲𝗽𝗼𝗿𝘁𝘀, 𝘀𝗰𝗮𝗻𝘀, 𝗼𝗿 𝗣𝗗𝗙 𝗽𝗮𝗴𝗲𝘀 𝗯𝘆 𝘁𝗵𝗲𝗶𝗿 𝘃𝗶𝘀𝘂𝗮𝗹 𝗰𝗼𝗻𝘁𝗲𝗻𝘁, 𝗺𝘂𝗹𝘁𝗶-𝘃𝗲𝗰𝘁𝗼𝗿 𝘄𝗶𝘁𝗵 𝗟𝗘𝗠𝗨𝗥 𝗼𝗿 𝗠𝗨𝗩𝗘𝗥𝗔 𝗶𝘀 𝗮 𝗺𝗲𝗮𝘀𝘂𝗿𝗮𝗯𝗹𝗲 𝘀𝘁𝗲𝗽 𝘂𝗽 𝗳𝗿𝗼𝗺 𝗮 𝗱𝗲𝗻𝘀𝗲 𝗯𝗮𝘀𝗲𝗹𝗶𝗻𝗲, 𝗮𝗻𝗱 𝘁𝗵𝗲 𝗹𝗲𝗮𝗱 𝘀𝘂𝗿𝘃𝗶𝘃𝗲𝘀 𝗼𝗻𝗰𝗲 𝘆𝗼𝘂 𝗺𝗼𝘃𝗲 𝗼𝗳𝗳 𝗲𝘅𝗮𝗰𝘁 𝘀𝗲𝗮𝗿𝗰𝗵. 💬 0 🔄 0 ❤️ 2 👀 85 ⚡