AI产品精选

Milvus 3.0 推出集合快照和 Spark 连接器,支持 AI 数据批量处理

𝗠𝗶𝗹𝘃𝘂𝘀 𝟯.𝟬: 𝗦𝘁𝗮𝗯𝗹𝗲 𝗦𝗻𝗮𝗽𝘀𝗵𝗼𝘁𝘀 𝗮𝗻𝗱 𝗕𝗮𝘁𝗰𝗵 𝗣𝗿𝗼𝗰𝗲𝘀𝘀𝗶𝗻𝗴 𝗳𝗼𝗿 ...

精选理由

Milvus 3.0 带快照和 Spark 连接器,不用复制数据就能给 AI 数据做稳定视图,还能跑分布式批量处理,把结果写回去。

AI 摘要

Milvus 3.0 引入 Collection Snapshot,通过引用现有数据文件、索引文件和元数据,在不复制完整数据集的情况下创建集合在特定时间点的只读视图。基于 Spark DataSource V2 的 Spark 连接器允许 Spark、Databricks 和 Amazon EMR 直接读写 Milvus。团队可用快照进行模型迁移、embedding 重新生成、schema 变更、回填、评估、去重和隔离测试。快照可作为逻辑恢复点,但不能替代备份,长期保留仍需独立备份。

原文 · Milvus

𝗠𝗶𝗹𝘃𝘂𝘀 𝟯.𝟬: 𝗦𝘁𝗮𝗯𝗹𝗲 𝗦𝗻𝗮𝗽𝘀𝗵𝗼𝘁𝘀 𝗮𝗻𝗱 𝗕𝗮𝘁𝗰𝗵 𝗣𝗿𝗼𝗰𝗲𝘀𝘀𝗶𝗻𝗴 𝗳𝗼𝗿 ...

𝗠𝗶𝗹𝘃𝘂𝘀 𝟯.𝟬: 𝗦𝘁𝗮𝗯𝗹𝗲 𝗦𝗻𝗮𝗽𝘀𝗵𝗼𝘁𝘀 𝗮𝗻𝗱 𝗕𝗮𝘁𝗰𝗵 𝗣𝗿𝗼𝗰𝗲𝘀𝘀𝗶𝗻𝗴 𝗳𝗼𝗿 𝗘𝘃𝗼𝗹𝘃𝗶𝗻𝗴 𝗔𝗜 𝗗𝗮𝘁𝗮 Milvus 3.0 introduces Collection Snapshot and a Spark connector to help teams create consistent views of changing AI data, run distributed batch processing, and write updated results back into Milvus. Traditional business data mainly records facts such as orders, product details, or device states. In AI systems, however, more data is generated by models. Embeddings, quality scores, classification labels, and reranking features all depend on specific models and processing logic. When models or business rules change, the same source data often needs to be recomputed, even while the production collection continues serving queries and accepting writes. 𝗖𝗼𝗹𝗹𝗲𝗰𝘁𝗶𝗼𝗻 𝗦𝗻𝗮𝗽𝘀𝗵𝗼𝘁 provides a stable input for these workflows. It creates a read-only view of a collection at a specific point in time by referencing existing data files, index files, and metadata, without copying the full dataset. Teams can use snapshots for model migration, embedding regeneration, schema changes, backfills, evaluation, deduplication, and isolated testing. Snapshots can also serve as logical recovery points, but they do not replace backups. A snapshot may still share the same underlying storage as production, making it better suited to short-term isolation and recovery. Independent backups remain necessary for long-term retention and disaster recovery. The Spark connector, built on Spark DataSource V2, allows Spark, Databricks, and Amazon EMR to read from and write to Milvus directly. Teams can process snapshot data with distributed jobs, then write new embeddings, labels, scores, or curated results back into Milvus. Snapshots provide a stable data view. Spark provides distributed compute. Milvus 3.0 connects them into one workf milvus.io/blog/announcin… ata improvement. More in the Milvus 3.0 launch blog: https://t.co/hq4qqZASuX 💬 0 🔄 0 ❤️ 0 👀 26 ⚡