2026年Embedding模型选择指南:从榜单到数据形态

Embedding model selection used to be mostly a leaderboard question. In 2026, it feels much closer t...

精选理由

Milvus团队写的实战指南,告诉你不同数据该用哪个embedding模型,还捎带介绍了Milvus 2.6的新功能。

AI 摘要

Milvus团队指出,2026年embedding模型选择更依赖数据形态而非排行榜。对于文本知识库,Qwen3 Embedding、Jina v5 text、BGE-M3、OpenAI text-embedding-3、Voyage、Cohere等密集embedding值得测试。对于PDF、图片、截图、音频、视频,推荐Gemini Embedding 2、Jina v5 omni、Cohere Embed 4、Qwen3-VL-Embedding、Voyage Multimodal等模型。长文档可用Voyage context-4进行上下文化片段嵌入。复杂布局可考虑ColPali风格的多向量检索。基准如MTEB、MMEB、ViDoRe可辅助筛选,但最终验证需用自己的文档和查询。Milvus 2.6新增Text Embedding Function,将嵌入管道移入数据库层。

原文 · Milvus

Embedding model selection used to be mostly a leaderboard question. In 2026, it feels much closer t...

Embedding model selection used to be mostly a leaderboard question. In 2026, it feels much closer to a data-shape question. A text knowledge base, a scanned contract, a product screenshot, a financial report, and a video archive contain different retrieval signals. Sending all of them through the same chunking and embedding pipeline can discard the information that matters most. 𝗙𝗼𝗿 𝗻𝗼𝗿𝗺𝗮𝗹 𝘁𝗲𝘅𝘁 𝗞𝗕𝘀, dense embeddings are still the practical starting point. Qwen3 Embedding, Jina v5 text, BGE-M3, OpenAI text-embedding-3, Voyage, and Cohere are all worth testing against your real chunks. 𝗙𝗼𝗿 𝗣𝗗𝗙𝘀, 𝗶𝗺𝗮𝗴𝗲𝘀, 𝘀𝗰𝗿𝗲𝗲𝗻𝘀𝗵𝗼𝘁𝘀, 𝗮𝘂𝗱𝗶𝗼, 𝗮𝗻𝗱 𝘃𝗶𝗱𝗲𝗼, models like Gemini Embedding 2, Jina v5 omni, Cohere Embed 4, Qwen3-VL-Embedding, and Voyage Multimodal are making it easier to keep more original context in retrieval. 𝗙𝗼𝗿 𝗹𝗼𝗻𝗴 𝗱𝗼𝗰𝘂𝗺𝗲𝗻𝘁𝘀, contextualized chunk embeddings like Voyage context-4 address a painful issue: chunks often lose meaning when separated from the full document. 𝗙𝗼𝗿 𝗰𝗼𝗺𝗽𝗹𝗲𝘅 𝗹𝗮𝘆𝗼𝘂𝘁𝘀, ColPali-style multi-vector retrieval is worth a look, especially when tables, figures, page regions, or scanned details decide the answer. Benchmarks like MTEB, MMEB, and ViDoRe are useful filters. But the final test still has to be your own documents, your own queries, and your own failure cases. 𝗠𝗶𝗹𝘃𝘂𝘀 𝟮.𝟲 also moves part of the embedding plumbing into the database layer: with Text Embedding Function, you can insert raw text, let Milvus call the configured embedding provider, store the vectors, and run text queries without managing embedding calls in every client. We wrote a practical guide on how to choose embedding models for the second half of 2026. 💬 0 🔄 0 ❤️ 0 👀 47 ⚡

2026年Embedding模型选择指南:从榜单到数据形态 · AI 热点