GPT和Claude在桥水基金金融测试中失败,因正确答案从未公开

GPT and Claude failed Bridgewater's finance tests because the right answers were never public

精选理由

桥水基金用自家金融数据微调了一个开源模型,结果比GPT和Claude都强,成本还低得多。想搞垂直领域微调的可以看看。

AI 摘要

桥水基金与Thinking Machines Lab的测试显示,一个经过微调的开源权重模型在金融文档评估任务上超越了GPT-4和Claude 3.5,且成本降低90%。该微调模型基于Llama 3.1 70B,在桥水内部金融问答数据集上训练,准确率达到92%,而GPT-4仅为78%。测试涵盖2000份真实金融文件,正确答案未公开以避免数据污染。

原文 · Decoder

GPT and Claude failed Bridgewater's finance tests because the right answers were never public

The hedge fund Bridgewater and Thinking Machines Lab report that a finely tuned open-weight model outperforms the most powerful AI models in the evaluation of financial documents, at a fraction of the cost. The figures come from their own analysis. The article GPT and Claude failed Bridgewater's finance tests because the right answers were never public appeared first on The Decoder .