精选理由
让用户直接看不同模型在同一任务上的输出对比,比只看基准数字直观多了,还能知道哪个实际表现更好。
Ethan Mollick在推文中建议AI实验室应推出易于观看的模型对比示例,比如展示Luna xHigh、Terra Medium、Sol Low在相同任务上的输出,或类似Haiku、Sonnet、Opus、Fable的对比。他认为需要展示错误率变化及用户在不同版本间的选择倾向,以帮助用户理解模型差异。
原文 · Ethan Mollick
The Labs should launch some easy-to-view compariso…
The Labs should launch some easy-to-view comparisons of their new models: show us what a piece of work looks like from Luna xHigh versus Terra Medium versus Sol Low (or similar for Haiku, Sonnet, Opus, Fable).
How do error rates change? Which do people pick when? Help is needed.