Perplexity开源内部研究基准WANDR
Perplexity把内部跑得最好的研究能力基准WANDR开源了,你要评测推理模型可以拿来用。
Perplexity宣布开源其内部用于衡量AI研究能力的基准WANDR。该基准在构建Perplexity Computer的深度研究功能中起到关键作用。公司表示WANDR在成本和性能上均达到最佳效果。Perplexity强调强大的内部评估和基准是其主要优势之一。
Perplexity has the best (both on cost and performance) deep and wide research harness in Computer. One of the contributing factors is strong internal evals and benchmarks. Today, we're open-sourcing WANDR, the benchmark we use internally for measuring research capabilities. Perplexity @perplexity_ai We’re open sourcing WANDR. WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer. research.perplexity.ai/articles/wandr… 🔗 View Quoted Tweet 💬 16 🔄 6 ❤️ 75 👀 7385 📊 19 ⚡