Milvus 用不到 1GB 内存跑 2500 万向量:FP16 + mmap 实战

𝗬𝗼𝘂 𝗰𝗮𝗻 𝗿𝘂𝗻 𝟮𝟱 𝗺𝗶𝗹𝗹𝗶𝗼𝗻 𝘃𝗲𝗰𝘁𝗼𝗿𝘀 𝗶𝗻 𝗠𝗶𝗹𝘃𝘂𝘀 𝘂𝘀𝗶𝗻𝗴 𝘂𝗻𝗱𝗲𝗿 ...

精选理由

做向量搜索的团队常被内存预算卡住,这个案例直接展示了 FLAT + FP16 + mmap 的组合拳如何把 139GB 需求压到 600MB,适合资源受限的单机部署场景,值得参考。

AI 摘要

Milvus 团队分享了一个用户案例:在单机 32GB 内存环境下,用 FLAT 索引配合 FP16 存储、mmap 内存映射和标量过滤,成功加载 2500 万 1280 维图像向量,实际驻留内存仅约 600MB,热查询延迟低于 100ms。默认 FP32 预估需 139GB,而 AISAQ 和 IVF_FLAT 索引均因构建或加载问题失败。该方案适合搜索空间远小于全量集合的场景,如租户级 RAG、带标签的图像搜索或电商搜索。

原文 · Milvus

𝗬𝗼𝘂 𝗰𝗮𝗻 𝗿𝘂𝗻 𝟮𝟱 𝗺𝗶𝗹𝗹𝗶𝗼𝗻 𝘃𝗲𝗰𝘁𝗼𝗿𝘀 𝗶𝗻 𝗠𝗶𝗹𝘃𝘂𝘀 𝘂𝘀𝗶𝗻𝗴 𝘂𝗻𝗱𝗲𝗿 ...

𝗬𝗼𝘂 𝗰𝗮𝗻 𝗿𝘂𝗻 𝟮𝟱 𝗺𝗶𝗹𝗹𝗶𝗼𝗻 𝘃𝗲𝗰𝘁𝗼𝗿𝘀 𝗶𝗻 𝗠𝗶𝗹𝘃𝘂𝘀 𝘂𝘀𝗶𝗻𝗴 𝘂𝗻𝗱𝗲𝗿 𝟭𝗚𝗕 𝗼𝗳 𝗺𝗲𝗺𝗼𝗿𝘆. A user had 25M image vectors, each with 1280 dimensions, and only 32GB of memory available for Milvus on a single machine. The default FP32 sizing estimate was 139GB. They first tried more advanced indexes, but neither worked out: • 𝗔𝗜𝗦𝗔𝗤 looked right for constrained hardware, but the build path was too heavy for the machine. • 𝗜𝗩𝗙_𝗙𝗟𝗔𝗧 built successfully, but the collection load hung at 14% and never finished. After working with our developers, the user switched to 𝗙𝗟𝗔𝗧, the simplest index in Milvus. FLAT avoided extra ANN structures and build/load complexity, while Milvus provided the pieces that made the setup practical: • 𝗙𝗣𝟭𝟲 storage cut each vector dimension from 4 bytes to 2 bytes, reducing raw vector data by half. • 𝗺𝗺𝗮𝗽 let Milvus access raw vector data through memory-mapped files instead of loading it all into process memory. • 𝗦𝗰𝗮𝗹𝗮𝗿 𝗳𝗶𝗹𝘁𝗲𝗿𝗶𝗻𝗴 narrowed each query first using fields like dataid and classid, so Milvus compared only a few thousand vectors instead of 25 million. 𝗧𝗵𝗲 𝗿𝗲𝘀𝘂𝗹𝘁: 𝗮𝗿𝗼𝘂𝗻𝗱 𝟲𝟬𝟬𝗠𝗕 𝗼𝗳 𝗿𝗲𝘀𝗶𝗱𝗲𝗻𝘁 𝗺𝗲𝗺𝗼𝗿𝘆 𝗮𝗻𝗱 𝘄𝗮𝗿𝗺 𝗾𝘂𝗲𝗿𝗶𝗲𝘀 𝘂𝗻𝗱𝗲𝗿 𝟭𝟬𝟬𝗺𝘀. When the real search space is much smaller than the full co milvus.io/blog/25-millio… enant RAG, labeled image search, or e-commerce search, FLAT + FP16 + mmap can be a practical option. Full breakdown in the blog: https://t.co/1aBXgMhf2l 💬 0 🔄 0 ❤️ 0 👀 24 ⚡