Gary Marcus一句话点破AI行业的真实痛点:基准测试再强,一到真实场景就露怯。适合关心AI落地效果的人。
Gary Marcus在推文中回应Derek Thompson关于‘模型强大为何世界如常’的疑问。他指出AI模型在基准测试(如GLUE、MMLU)中得分虽高,但在现实世界可靠性不足,难以对多数企业产生变革性影响。Marcus认为不存在真正的悖论,而是基准测试与实用价值之间仍存在显著鸿沟。
there is no paradox here. the models may be “powerful” as measured by benchmarks but in the real wo...
there is no paradox here. the models may be “powerful” as measured by benchmarks but in the real world they still just aren’t reliable enough be all that transformative for most businesses. Derek Thompson @DKThomp I think a reasonably fair question to ask AI ppl these days is something like, “If the models are so powerful, why does the world seem so normal?” 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 2 👀 294 📊 1 ⚡