Milvus 3.0 新出的 External Collections,让你直接搜对象存储里的向量数据,不用再拷贝一份。省掉同步和权限麻烦,快试试。
Milvus 3.0 推出 External Collections 功能,允许用户直接对存储在 S3、Parquet 文件、Lance 数据集或 Iceberg 表等对象存储中的向量数据进行搜索,无需复制到服务数据库。此前有两种方案:复制数据到向量数据库(低延迟但需维护同步)或直接查询湖(无索引全量扫描)。External Collections 提供第三种路径:在源数据上构建向量、BM25 倒排、JSON 和标量索引,数据不移动。该功能为只读零拷贝,支持增量索引新片段,并提供三种加载模型以平衡存储成本与延迟。适用于需要生产级搜索但避免数据复制的湖数据集。
⭐ 𝗠𝗶𝗹𝘃𝘂𝘀 𝟯.𝟬 𝗶𝗻𝘁𝗿𝗼𝗱𝘂𝗰𝗲𝘀 𝗘𝘅𝘁𝗲𝗿𝗻𝗮𝗹 𝗖𝗼𝗹𝗹𝗲𝗰𝘁𝗶𝗼𝗻𝘀, 𝗮 𝘄𝗮𝘆 𝘁𝗼 ...
⭐ 𝗠𝗶𝗹𝘃𝘂𝘀 𝟯.𝟬 𝗶𝗻𝘁𝗿𝗼𝗱𝘂𝗰𝗲𝘀 𝗘𝘅𝘁𝗲𝗿𝗻𝗮𝗹 𝗖𝗼𝗹𝗹𝗲𝗰𝘁𝗶𝗼𝗻𝘀, 𝗮 𝘄𝗮𝘆 𝘁𝗼 𝗺𝗮𝗸𝗲 𝗹𝗮𝗸𝗲-𝗿𝗲𝘀𝗶𝗱𝗲𝗻𝘁 𝘃𝗲𝗰𝘁𝗼𝗿 𝗱𝗮𝘁𝗮 𝘀𝗲𝗮𝗿𝗰𝗵𝗮𝗯𝗹𝗲 𝘄𝗶𝘁𝗵𝗼𝘂𝘁 𝗰𝗼𝗽𝘆𝗶𝗻𝗴 𝗶𝘁 𝗶𝗻𝘁𝗼 𝗮 𝘀𝗲𝗿𝘃𝗶𝗻𝗴 𝗱𝗮𝘁𝗮𝗯𝗮𝘀𝗲. Many teams already have embeddings and metadata in object storage: Parquet files in S3, Lance datasets, Iceberg tables, or other lakehouse formats. Before Milvus 3.0, there were usually two ways to make that data searchable. 𝗢𝗽𝘁𝗶𝗼𝗻 𝗼𝗻𝗲: 𝗰𝗼𝗽𝘆 𝗶𝘁 𝗶𝗻𝘁𝗼 𝗮 𝘃𝗲𝗰𝘁𝗼𝗿 𝗱𝗮𝘁𝗮𝗯𝗮𝘀𝗲. You get low-latency ANN search, but now you have a second copy and an ETL pipeline to keep in sync. 𝗢𝗽𝘁𝗶𝗼𝗻 𝘁𝘄𝗼: 𝗾𝘂𝗲𝗿𝘆 𝘁𝗵𝗲 𝗹𝗮𝗸𝗲 𝗱𝗶𝗿𝗲𝗰𝘁𝗹𝘆. You avoid duplication, but without ANN indexes, vector search turns into a brute-force scan. 𝗘𝘅𝘁𝗲𝗿𝗻𝗮𝗹 𝗖𝗼𝗹𝗹𝗲𝗰𝘁𝗶𝗼𝗻𝘀 𝗶𝗻𝘁𝗿𝗼𝗱𝘂𝗰𝗲 𝗮 𝘁𝗵𝗶𝗿𝗱 𝗽𝗮𝘁𝗵. You keep the data where it is, map external fields into a Milvus schema, and use the same Milvus search and query APIs. Milvus builds vector, BM25 inverted, JSON, and scalar indexes over the lake-resident data. The source files do not move. For teams where the lake owns permissions and freshness, every extra copy creates sync, access-control, and debugging work. 𝗔 𝗳𝗲𝘄 𝗽𝗿𝗮𝗰𝘁𝗶𝗰𝗮𝗹 𝗱𝗲𝘁𝗮𝗶𝗹𝘀 𝗺𝗮𝘁𝘁𝗲𝗿: • External Collections are read-only and zero-copy. • Milvus can index newly added fragments instead of rebuilding the whole collection. • Three load m milvus.io/docs/create-an… etween lower storage cost and lower latency. Native Milvus collections are better for write-heavy serving. External Collections are for lake datasets that need production search without another copy. Know the details: https://t.co/NR2QfOORyC 💬 0 🔄 0 ❤️ 0 👀 55 ⚡