模型精选

腾讯发布探索能力基准测试ExplorationBench

more 🔛 https://t.co/XsFS9gaRpQ

精选理由

腾讯团队出了个新基准测试,专门测AI的探索能力,结果挺有意思的。

腾讯、复旦大学和清华大学联合推出探索能力基准测试ExplorationBench,包含两个沙箱环境AlienCode和AlienLogic。测试覆盖10个前沿AI系统,结果显示获取反馈比独自思考更有效,实验设计对性能影响显著。即使规则正确陈述,任务完成率也只有73.4%,同一系统在不同环境下的得分差异巨大。

图片来源 · Hunyuan
原文 · Hunyuan

more 🔛 https://t.co/XsFS9gaRpQ

more 🔛 github.com/Tencent-Hunyua… Q Tencent Hy @TencentHunyuan New Research: We are releasing ExplorationBench, a benchmark for measuring how AI systems explore. Scientific discovery begins where known problems end: a system has to frame hypotheses, design experiments, and learn from the results. Evaluating this is hard. Genuinely new answers cannot be checked quickly, and in familiar domains a model can simply recall what it has seen. Addressing this challenge, researchers from Tencent Hy, Fudan University, and Tsinghua University built verifiable Alien Worlds. Their rules are executable, so every answer is checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. 🔹 Two sandboxes: AlienCode (31 hidden rule changes, 70 tasks) and AlienLogic (24 patched inference rules, 70 theorems) 🔹 A flawed manual, four rounds of self-designed probes, and closed-book tests after every round 🔹 Every answer graded by an interpreter or a proof checker, with no LLM judge What we found across 10 frontier AI systems: 1️⃣ Getting feedback is more effective than thinking alone. No AlienCode run starts above 15.7%; after four rounds the best reaches 89.0%, while the same turns without feedback stay at 0.5–11.0%. 2️⃣ Designing the experiments matters. Replaying a system's own best probes gives it exactly the same evidence, yet in AlienCode 9 of 10 systems do worse than when they chose the probes themselves. 3️⃣ Knowing a rule is not using it. Even when every required rule is stated correctly, tasks are solved only 73.4% of the time. 4️⃣ One score hides a lot. The same system under the same budget ended anywhere from 5.7% to 79.0%, and rankings barely transfer between the two worlds. CL-bench asked whether models can learn from context. ExplorationBench asks whether they can discover the rules themselves. 📄 Pap arxiv.org/abs/2609.30199 JyxQ 🌐 Website & leaderbo explorationbench.com Cfc2s 📝 explorationbench.com/blog/ H7ZnQH 💻 Code (coming github.com/Tencent-Hunyua… lgErkcl 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 155 ⚡