Milvus 团队用三个开源项目实测 Jev:便宜快速但上限明显
𝗪𝗲 𝗶𝗻𝘁𝗲𝗴𝗿𝗮𝘁𝗲𝗱 𝗝𝗲𝘃 𝗶𝗻𝘁𝗼 𝘁𝗵𝗿𝗲𝗲 𝗼𝗽𝗲𝗻-𝘀𝗼𝘂𝗿𝗰𝗲 𝗽𝗿𝗼𝗷𝗲𝗰𝘁𝘀 𝘁𝗼 ...
Zilliz 团队拿 Jev 跑了三个真实场景,省到七分之一的成本,性能掉多少都给你列出来了,做 RAG 的可以抄作业。
Milvus 团队把 Jev 集成到 DeepSearcher、MemSearch、Vector Graph RAG 三个开源项目中做评测。在 DeepSearcher 的 100 个多跳问题上,Jev 与 DeepSeek V4 Flash 同样达到 93.25% Recall @5,但决策延迟中位数从 2.23s 降到 0.55s,API 成本约为七分之一。在 MemSearch 的 2,172 条中英文查询上,Jev 把 Recall @5 从 74.71% 提升到 79.41%,仍落后 Voyage rerank-3 的 81.87%。在 Vector Graph RAG 中,Jev 超过 GPT-4o-mini,但在 HotpotQA 和 MuSiQue 上分别落后 GPT-5-mini 1 分和 4.13 分。结论是 Jev 适合搜索停止、过滤、路由等有边界的频繁判断任务,推理密集场景仍需更强的生成模型。
𝗪𝗲 𝗶𝗻𝘁𝗲𝗴𝗿𝗮𝘁𝗲𝗱 𝗝𝗲𝘃 𝗶𝗻𝘁𝗼 𝘁𝗵𝗿𝗲𝗲 𝗼𝗽𝗲𝗻-𝘀𝗼𝘂𝗿𝗰𝗲 𝗽𝗿𝗼𝗷𝗲𝗰𝘁𝘀 𝘁𝗼 ...
𝗪𝗲 𝗶𝗻𝘁𝗲𝗴𝗿𝗮𝘁𝗲𝗱 𝗝𝗲𝘃 𝗶𝗻𝘁𝗼 𝘁𝗵𝗿𝗲𝗲 𝗼𝗽𝗲𝗻-𝘀𝗼𝘂𝗿𝗰𝗲 𝗽𝗿𝗼𝗷𝗲𝗰𝘁𝘀 𝘁𝗼 𝘁𝗲𝘀𝘁 𝘁𝗵𝗮𝘁 𝗶𝗱𝗲𝗮. First, in 𝗗𝗲𝗲𝗽𝗦𝗲𝗮𝗿𝗰𝗵𝗲𝗿, we used Jev to decide whether an Agentic Search workflow had gathered enough evidence to stop. Across 100 multi-hop questions, Jev matched DeepSeek V4 Flash at 93.25% Recall @5 . Yet median decision latency fell from 2.23s to 0.55s, with estimated API cost at roughly one-seventh. Second, in 𝗠𝗲𝗺𝗦𝗲𝗮𝗿𝗰𝗵, we tested coding-agent memory reranking on 2,172 private Chinese and English queries. Jev improved Recall @5 from 74.71% to 79.41%, but remained behind the specialist Voyage rerank-3 at 81.87%. Third, in 𝗩𝗲𝗰𝘁𝗼𝗿 𝗚𝗿𝗮𝗽𝗵 𝗥𝗔𝗚, Jev reranked and filtered candidate relationships for multi-hop retrieval. It outperformed GPT-4o-mini, but trailed GPT-5-mini—by just 1 point on HotpotQA and 4.13 points on the harder MuSiQue benchmark. Jev is compelling for frequent, bounded judgments such as search stopping, filtering, routing, and simpler reranking. But as tasks become more domain-specific or reasoning-heavy, stronger generative models and specialist rerankers still earn their place. Want to di github.com/zilliztech/dee… • DeepSearcher — search-stopping experiment, code, and e github.com/zilliztech/mem… /aqPCKQnhMi • MemSearch — Jev reranking integration and evaluation: https://t.co/ github.com/zilliztech/vec… ph RAG — Jev relationship filtering, evaluation, and cost analysis: https://t.co/7z9kGZsKuO 💬 0 🔄 0 ❤️ 0 👀 37 ⚡