技巧精选

RAG管道先嵌入还是先分块?Max-Min语义分块解析

𝗙𝗼𝗿 𝗮 𝗥𝗔𝗚 𝗽𝗶𝗽𝗲𝗹𝗶𝗻𝗲, 𝗱𝗼 𝘆𝗼𝘂 𝗰𝗵𝘂𝗻𝗸 𝗳𝗶𝗿𝘀𝘁 𝗼𝗿 𝗲𝗺𝗯𝗲𝗱 𝗳𝗶𝗿𝘀𝘁? The...

精选理由

Milvus团队介绍的Max-Min语义分块方法先嵌入句子再依据语义决定分块边界,比传统固定长度分块更聪明,适合想提升RAG检索质量的人学习。

AI 摘要

标准RAG管道采用先分块后嵌入的流程:分块→嵌入→索引→检索→生成,但固定长度或递归分割的分块方式存在精度与上下文的权衡。Max-Min Semantic Chunking采用先嵌入后分块模式:先用文本嵌入模型将所有句子映射到高维空间,再通过计算块内最小余弦相似度和新句子与块的最大余弦相似度决定是否合并。该方法的决策规则是:若新句子与块的最强连接大于块内最弱内部链接,则加入当前块,否则开启新块。它维护句子顺序,通过自适应大小限制和相似度阈值保持块连贯性。但因为它按顺序局部聚类,可能遗漏长文档中跨段落的远距离依赖。

原文 · Milvus

𝗙𝗼𝗿 𝗮 𝗥𝗔𝗚 𝗽𝗶𝗽𝗲𝗹𝗶𝗻𝗲, 𝗱𝗼 𝘆𝗼𝘂 𝗰𝗵𝘂𝗻𝗸 𝗳𝗶𝗿𝘀𝘁 𝗼𝗿 𝗲𝗺𝗯𝗲𝗱 𝗳𝗶𝗿𝘀𝘁? The...

𝗙𝗼𝗿 𝗮 𝗥𝗔𝗚 𝗽𝗶𝗽𝗲𝗹𝗶𝗻𝗲, 𝗱𝗼 𝘆𝗼𝘂 𝗰𝗵𝘂𝗻𝗸 𝗳𝗶𝗿𝘀𝘁 𝗼𝗿 𝗲𝗺𝗯𝗲𝗱 𝗳𝗶𝗿𝘀𝘁? The standard RAG pipeline chunks documents before embedding them. 𝗔 𝗱𝗶𝗳𝗳𝗲𝗿𝗲𝗻𝘁 𝗽𝗮𝘁𝘁𝗲𝗿𝗻 𝗶𝘀 𝘀𝘁𝗮𝗿𝘁𝗶𝗻𝗴 𝘁𝗼 𝘀𝗵𝗼𝘄 𝘂𝗽: 𝗲𝗺𝗯𝗲𝗱 𝗮𝗹𝗹 𝘀𝗲𝗻𝘁𝗲𝗻𝗰𝗲𝘀 𝗳𝗶𝗿𝘀𝘁, 𝘁𝗵𝗲𝗻 𝘂𝘀𝗲 𝘁𝗵𝗲𝗶𝗿 𝘀𝗲𝗺𝗮𝗻𝘁𝗶𝗰 𝘀𝗶𝗺𝗶𝗹𝗮𝗿𝗶𝘁𝘆 𝘁𝗼 𝗱𝗲𝗰𝗶𝗱𝗲 𝘄𝗵𝗲𝗿𝗲 𝘁𝗼 𝗰𝗵𝘂𝗻𝗸. Max-Min Semantic Chunking is one example. 𝗪𝗵𝘆 𝗲𝗺𝗯𝗲𝗱 𝗳𝗶𝗿𝘀𝘁? Because the standard RAG pipeline chunks documents before embedding them, which means boundaries are usually drawn without embedding-based similarity signals. Embedding all sentences upfront gives the algorithm similarity data to work with, so boundaries can follow actual shifts in meaning. 𝗜𝗻 𝗺𝗮𝗻𝘆 𝗥𝗔𝗚 𝗽𝗶𝗽𝗲𝗹𝗶𝗻𝗲𝘀, 𝗰𝗵𝘂𝗻𝗸𝗶𝗻𝗴 𝗵𝗮𝗽𝗽𝗲𝗻𝘀 𝗯𝗲𝗳𝗼𝗿𝗲 𝗲𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴, 𝗮𝗻𝗱 𝗰𝗵𝘂𝗻𝗸 𝗾𝘂𝗮𝗹𝗶𝘁𝘆 𝘀𝘁𝗿𝗼𝗻𝗴𝗹𝘆 𝗮𝗳𝗳𝗲𝗰𝘁𝘀 𝗿𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹 𝗾𝘂𝗮𝗹𝗶𝘁𝘆. The conventional sequence runs: chunk → embed → index in a vector database like Milvus → retrieve → generate. But at the chunking step, fixed-length and recursive splitting do not fully remove the tradeoff. Smaller chunks tend to improve precision but can lose context. Larger chunks preserve more context but can introduce noise. 𝗠𝗮𝘅-𝗠𝗶𝗻 𝗦𝗲𝗺𝗮𝗻𝘁𝗶𝗰 𝗖𝗵𝘂𝗻𝗸𝗶𝗻𝗴 𝗶𝘀 𝗼𝗻𝗲 𝗺𝗲𝘁𝗵𝗼𝗱 𝗴𝗮𝗶𝗻𝗶𝗻𝗴 𝘁𝗿𝗮𝗰𝘁𝗶𝗼𝗻 𝘄𝗶𝘁𝗵 𝘁𝗵𝗲 𝗲𝗺𝗯𝗲𝗱-𝗳𝗶𝗿𝘀𝘁 𝗮𝗽𝗽𝗿𝗼𝗮𝗰𝗵. It treats chunking as a sequential clustering problem: walk through the document in order, decide at each sentence whether it joins the current chunk or starts a new one. 𝗦𝗲𝗻𝘁𝗲𝗻𝗰𝗲𝘀 𝗺𝘂𝘀𝘁 𝘀𝘁𝗮𝘆 𝗰𝗼𝗻𝘀𝗲𝗰𝘂𝘁𝗶𝘃𝗲, 𝗮𝗻𝗱 𝘁𝗵𝗲 𝗮𝗹𝗴𝗼𝗿𝗶𝘁𝗵𝗺 𝗻𝗲𝘃𝗲𝗿 𝗿𝗲𝗮𝗿𝗿𝗮𝗻𝗴𝗲𝘀 𝘁𝗲𝘅𝘁. 𝗛𝗲𝗿𝗲'𝘀 𝗵𝗼𝘄 𝗶𝘁 𝘄𝗼𝗿𝗸𝘀, 𝘀𝘁𝗲𝗽 𝗯𝘆 𝘀𝘁𝗲𝗽: • 𝗘𝗺𝗯𝗲𝗱 𝘁𝗵𝗲 𝘀𝗲𝗻𝘁𝗲𝗻𝗰𝗲𝘀 𝗳𝗶𝗿𝘀𝘁: A text embedding model maps all sentences to high-dimensional space. Suppose the first n−k sentences have been assigned to the current chunk C. The decision now is whether sentence n−k+1 joins C or starts a new chunk. • 𝗖𝗼𝗺𝗽𝘂𝘁𝗲 𝗺𝗶𝗻𝗶𝗺𝘂𝗺 𝘀𝗶𝗺𝗶𝗹𝗮𝗿𝗶𝘁𝘆 𝘄𝗶𝘁𝗵𝗶𝗻 𝘁𝗵𝗲 𝗰𝗵𝘂𝗻𝗸: Calculate the minimum pairwise cosine similarity among all sentence vectors in C. This identifies the most semantically dissimilar pair, measuring how tightly the group holds together, and sets the bar for whether the new sentence belongs. • 𝗖𝗼𝗺𝗽𝘂𝘁𝗲 𝗺𝗮𝘅𝗶𝗺𝘂𝗺 𝘀𝗶𝗺𝗶𝗹𝗮𝗿𝗶𝘁𝘆 𝘁𝗼 𝘁𝗵𝗲 𝗻𝗲𝘄 𝘀𝗲𝗻𝘁𝗲𝗻𝗰𝗲: Calculate the maximum cosine similarity between the new sentence and every sentence in C. This captures the strongest connection the incoming sentence has to the existing group. • 𝗧𝗵𝗲 𝗱𝗲𝗰𝗶𝘀𝗶𝗼𝗻 𝗿𝘂𝗹𝗲: If the new sentence's strongest connection to the chunk beats the chunk's weakest internal link, it joins. Otherwise, a new chunk begins. • 𝗨𝘀𝗲 𝘀𝗶𝘇𝗲 𝗹𝗶𝗺𝗶𝘁𝘀 𝗮𝗻𝗱 𝘀𝗶𝗺𝗶𝗹𝗮𝗿𝗶𝘁𝘆 𝘁𝗵𝗿𝗲𝘀𝗵𝗼𝗹𝗱𝘀 𝘁𝗼 𝗸𝗲𝗲𝗽 𝗰𝗵𝘂𝗻𝗸𝘀 𝗳𝗿𝗼𝗺 𝗴𝗿𝗼𝘄𝗶𝗻𝗴 𝘁𝗼𝗼 𝗹𝗼𝗼𝘀 milvus.io/blog/embedding… unk coherence, dynamically adjust chunk size limits and similarity thresholds. • 𝗜𝗻𝗶𝘁𝗶𝗮𝗹𝗶𝘇𝗮𝘁𝗶𝗼𝗻: When the chunk has only one sentence, there's no internal minimum to compute. Compare the similarity between the first and second sentence against a preset threshold constant. Above it, they chunk together. Below it, they split. 𝗠𝗮𝘅-𝗠𝗶𝗻 𝗵𝗮𝘀 𝗶𝘁𝘀 𝘁𝗿𝗮𝗱𝗲𝗼𝗳𝗳: because it clusters sequentially and locally, it can miss long-range context dependencies in long documents. Important information spread across distant sections may end up in separate chunks. Full walkthrough with code: https://t.co/QqyHQ4U2Rq 💬 0 🔄 0 ❤️ 0 👀 31 ⚡

RAG管道先嵌入还是先分块?Max-Min语义分块解析 · AI 热点