Check out Milvus 3.0's regex capabilities for efficient row scanning. It's a game-changer for handling large datasets with RE2's speed and Milvus' optimizations compared to other regex solutions.
Milvus 3.0 introduces regex support for safe and predictable scanning of millions of rows. RE2 provides linear-time matching, and Milvus optimizes regex operations with various techniques. Regex filtering is now available in Milvus 3.0.
Regex becomes an engineering problem when it has to scan millions of rows safely and predictably. ...
Regex becomes an engineering problem when it has to scan millions of rows safely and predictably. 𝗠𝗶𝗹𝘃𝘂𝘀 𝟯.𝟬 𝗮𝗱𝗱𝘀 𝘁𝗵𝗲 =~ 𝗮𝗻𝗱 !~ 𝗳𝗶𝗹𝘁𝗲𝗿𝘀 for matching or excluding structural string patterns alongside vector search, full-text search, and scalar filters. 𝗨𝗻𝗱𝗲𝗿 𝘁𝗵𝗲 𝗵𝗼𝗼𝗱, 𝗠𝗶𝗹𝘃𝘂𝘀 𝘂𝘀𝗲𝘀 𝗮 𝗹𝗮𝘆𝗲𝗿𝗲𝗱 𝗲𝘅𝗲𝗰𝘂𝘁𝗶𝗼𝗻 𝗽𝗮𝘁𝗵: 1. 𝗩𝗮𝗹𝗶𝗱𝗮𝘁𝗲 𝘄𝗶𝘁𝗵 𝗥𝗘𝟮. RE2 provides linear-time matching, and invalid patterns are rejected before the scan begins. 2. 𝗥𝗲𝘄𝗿𝗶𝘁𝗲 𝘀𝗶𝗺𝗽𝗹𝗲 𝗽𝗮𝘁𝘁𝗲𝗿𝗻𝘀. ^ERROR$ can become equality, while anchored literals can become prefix or suffix checks. The planner also evaluates cheaper predicates before regex. 3. 𝗥𝗲𝗱𝘂𝗰𝗲 𝘄𝗼𝗿𝗸 𝗼𝗻 𝗿𝗮𝘄 𝘀𝗰𝗮𝗻𝘀. On sealed segments, Milvus reuses a compiled pattern and checks required literals first, reducing the number of strings that need full RE2 verification. 4. 𝗨𝘀𝗲 𝗡𝗚𝗥𝗔𝗠 𝘄𝗵𝗲𝗻 𝗶𝘁 𝗰𝗮𝗻 𝗺𝗲𝗮𝗻𝗶𝗻𝗴𝗳𝘂𝗹𝗹𝘆 𝗿𝗲𝗱𝘂𝗰𝗲 𝗰𝗮𝗻𝗱𝗶𝗱𝗮𝘁𝗲𝘀. Milvus extracts required literals, intersects their posting lists, and runs RE2 only on the surviving strings. If it cannot safely narrow the candidate set, it falls back to raw verification. That qualification matters: NGRAM works best for literal-rich patterns with low candidate rates. It may offer little benefit when most rows survive the first phase. One more detail: w milvus.io/blog/milvus-3-… missing JSON paths remain UNKNOWN. Negation does not turn missing data into a match. Regex filtering is available in Milvus 3.0. Give it a try: https://t.co/TQE8EOUaB0 💬 0 🔄 0 ❤️ 0 👀 58 ⚡