想精确知道用AI干活到底省不省钱?METR直接算了一笔账,结果初期表现还不咋样。
METR(AI安全研究组织)提出了名为“expenditure horizon”的新指标,用于量化AI智能体在解决问题时的成本效益。在NanoGPT speedrun任务上的早期测试显示,AI智能体的成本效益低于人类。该指标存在盲点,例如未计算模型训练成本。随着新一代推理模型的发展,这一评估结果可能被改写。
METR introduces a new metric to calculate exactly when AI agents become more expensive than humans
METR's new metric, the "expenditure horizon," puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming, the metric has blind spots, and the newest generation of models could change the picture. The article METR introduces a new metric to calculate exactly when AI agents become more expensive than humans appeared first on The Decoder .