Milvus 3.0 Spark Connector发布

Your Spark job has finished generating embedding_v2. Now you need to add it to hundreds of millions ...

精选理由

Milvus 3.0 Spark Connector解决了大规模向量数据更新的难题,支持离线处理与实时写入分离。

AI 摘要

Milvus 3.0 Spark Connector支持将embedding_v2添加到数亿现有记录而不影响实时写入。该连接器通过快照固定模式、段布局和存储元数据,将Parquet输出与现有行通过主键连接。每个段成为独立执行单元,支持验证、重试和审计,并携带模式版本以拒绝过期输出。

原文 · Milvus

Your Spark job has finished generating embedding_v2. Now you need to add it to hundreds of millions ...

Your Spark job has finished generating embedding_v2. Now you need to add it to hundreds of millions of existing records in Milvus—without slowing down live writes. Sending the Parquet output back as millions of upserts forces the offline job to compete with production traffic for network, compute, and write capacity. If it fails halfway, tracking and retrying the unfinished work is another problem. 𝗕𝗮𝗰𝗸𝗳𝗶𝗹𝗹 𝗶𝗻 𝘁𝗵𝗲 𝗠𝗶𝗹𝘃𝘂𝘀 𝟯.𝟬 𝗦𝗽𝗮𝗿𝗸 𝗖𝗼𝗻𝗻𝗲𝗰𝘁𝗼𝗿 𝗸𝗲𝗲𝗽𝘀 𝘁𝗵𝗮𝘁 𝘄𝗼𝗿𝗸 𝗼𝗳𝗳𝗹𝗶𝗻𝗲: → A snapshot pins the schema, segment layout, and storage metadata. → Spark joins the Parquet output to existing rows by primary key. → Only the target columns are rebuilt, while row alignment is preserved. → Each segment becomes an independent unit for execution, validation, retry, and audit. The result also carries its schema version, so stale output can be rejected before submission. The goal isn’t faster upserts. It’s a controlled path for publishing embedding_v2 without turning the produc milvus.io/blog/announcin… e migration engine. Learn more: https://t.co/uZxGB2kxuk 💬 0 🔄 0 ❤️ 0 👀 85 ⚡