论文精选

模型评估讨论:测试套件问题

Important discussion. Measuring models against harnesses is completely broken. I prefer to test mode...

精选理由

Omar Solmaz讨论了模型评估的问题,建议关注标准测试套件的使用,了解模型在不同套件下的表现差异。

AI 摘要

模型评估方式存在问题,建议使用标准测试套件。模型质量应与最小套件如Pi和Hermes Agent比较。目前缺乏标准化方法,但模型可能未来能动态生成套件。Claude模型在这方面已有尝试,但不够一致。测试套件工程是AI公司关注的焦点,但优化过于个性化。

原文 · elvis

Important discussion. Measuring models against harnesses is completely broken. I prefer to test mode...

Important discussion. Measuring models against harnesses is completely broken. I prefer to test model quality against minimal harnesses like Pi and Hermes Agent. This is not perfect, as there are biases in the harnesses that favor some models and not others. A standard way to do this is missing but important, as harness engineering is where leading AI companies are focusing efforts. Not enough effort here, as things are moving fast and harness optimization is individualistic. On the flip side, I feel like models will eventually have the ability to dynamically generate harnesses on the fly as per task. Claude models do this already to some extent, though pretty inconsistently and remain a mystery. But this could mean that a harness is just a tunable artifact like a system prompt. In that realm, how are we assessing it, and exactly what? Benchmarking will only get murkier from here onwards. Onur Solmaz @onusoz We need to normalize measuring and judging models against a standardized test harness "Oh but model X performs best in their own proprietary harness" I could not care less. When I take exams, I go to the standardized classroom, get the standardized pencil and exam sheet, and have to solve it under 2 hours This system arose because we have a LOT of people to test Guess what? We now have a LOT of models, and they are multiplying by the day "Oh but model X performs substantially better in ARC-AGI-3 with a custom harness" I don't care... Then imbue model X with enough knowledge so that it can reconstruct that harness on the spot The main harness could be mini-swe-agent, terminus 2, vanilla pi or something along those lines It needs to be simple, and stay roughly the same over time There is already too much complexity in the benchmarking space right now, and I feel like not enough people are putting their feet down to cut away some part of it 🔗 View Quoted Tweet 💬 4 🔄 1 ❤️ 11 👀 1964 📊 4 ⚡