AI评估工具使用体验分享
HamelHusain实测了AI评估工具,能自动发现多种问题但流程有缺陷,适合想了解评估工具优缺点的开发者。
开发者HamelHusain测试了AI评估工具,发现其存在先创建评估再查看数据的问题。该工具能自动发现人类交接、格式化、语音助手等多种问题,但评估范围过于宽泛。相比之前的评估插件,用户体验有明显提升,但建议用户先查看数据再创建自己的标注应用。
Put this through its paces yesterday. Some thoughts:
The Bad:
1. The workflow creates evals before looking at data, and has you correct its mistakes. I believe you should be looking at data first to inform your understanding and to prioritize what to work on before trying to write evals.
2. It created markdown files for looking at data and asked us to "tell it" which labels were off. You should create your own annotation app instead. You are using coding agents after all (we did this in our livestream)!
3. The evaluators it scoped were too broad, and bundled too many types of failures at once. This can be avoided with looking at the data first. I'm afraid this might steer people in the wrong direction.
4. The workflow put us DEEP into the rabbit hole of a specific eval right away. It also asks "does it look right" at several steps without really giving you enough information to make that determination. I think this will steer people towards the wrong workflow in many cases.
The Good:
1. I was impressed by out of the box ability for issue discovery that other auto-eval approaches haven't been able to find! It found issues with human handoff, formatting, voice agents, and more. It's still better to look at your data iteratively with an agent, but this was the strongest performance I've seen with a more "one-shot" issue discovery approach.
2. This is a big improvement in terms of UX from their prior eval plugin thing https://t.co/oae9HkUqvz
3. The blog post they released conveys thinking that I agree with, such as the importance of looking at data, sampling intelligently, not saturating your own evals, etc. I'm really happy more people are thinking about evals this way.
Video of our attempt at using this here: https://t.co/mIqOonJAFb