MRVQ:单一索引同时支持维度与码率弹性向量检索
MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search
搞向量检索的可以看这篇:一份 MRVQ 索引同时搞定降维和降码率,内存省到 20 倍,还老实交代了质量上的取舍。
arXiv 论文提出 Matryoshka Residual Vector Quantization(MRVQ),一种面向冻结嵌入的事后残差量化器,通过丢弃残差阶段降码率、丢弃嵌入坐标降维度,让一份常驻索引覆盖所有(维度,码率)组合。在 FiQA 和 NFCorpus 上测试 4 个嵌入家族和 4/8/16 字节编码,MRVQ 内存比分别训练的三个 QINCo2 索引低 17.8-22.0 倍,比精简共享模型低 1.89-2.02 倍。代价是按码率单独训练的 QINCo2 在 FiQA 上 nDCG@10 高 0.026-0.107,但 MRVQ 在同等码长下优于 PQ、OPQ 和 AdANNS-OPQ。论文还给出 PCA-scalar 低成本设计,构建中位数速度快 420 倍,质量接近 RaBitQ,并报告了两项负面结果。
MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search
Dense-retrieval services must switch among embedding-prefix dimensions and index bit rates as latency, quality, and memory budgets change. Tuning a quantizer separately for each rate gives the best quality, but the retrieval tier then holds several code streams and quantizer states at once. We introduce Matryoshka Residual Vector Quantization (MRVQ), a post-hoc residual quantizer for frozen embeddings. Its maximum-rate code can be truncated two ways: dropping residual stages lowers the rate, and dropping embedding coordinates lowers the dimension. One resident artifact therefore serves every (dimension, rate) pair we evaluate. Across FiQA and NFCorpus, four embedding families, and {4, 8, 16}-byte codes, MRVQ is the lowest-RAM design we evaluate. It uses 17.8-22.0x less memory than three separately trained QINCo2 indices, and 1.89-2.02x less than a lean shared-model steelman. The saving is not free: per-rate QINCo2 is 0.026-0.107 nDCG@10 better on FiQA. But MRVQ beats PQ, OPQ, and AdANNS-OPQ at matched code size. We also evaluate a low-build-cost PCA-scalar design that attains quality comparable to RaBitQ and its extension while fitting 420x faster at the median. Finally, we report two negative results: QINCo2 collapses when trained at high rates, and a ranking-bound hypothesis misses its pre-specified acceptance criteria. MRVQ is therefore a low-memory operating point for elastic retrieval, not a universal quality winner.