模型

Perplexity 用提示引导自蒸馏训练 Computer 模型从错误中学习

Super cool!

精选理由

Perplexity 分享了怎么让 Computer 模型从自己的错误里学习,A/B 测试里工具调用失败率降了 21.2%,做法挺值得看看的。

Perplexity 公布了一项新研究,他们采用 hint-guided self-distillation(提示引导自蒸馏)方法对 Computer 模型做后训练,让模型从自身错误中学习。在一次线上 A/B 测试中,较晚训练的 checkpoint 相比较早的 checkpoint 将工具调用失败率相对降低了 21.2%。

原文 · Suhail

Super cool!

Super cool! Perplexity @perplexity_ai New research: We post-trained a Computer model to learn from its own errors using hint-guided self-distillation. In a live A/B test, a later trained checkpoint reduced tool-call failures by 21.2% relative to an earlier checkpoint. 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 4 👀 2054 📊 1 ⚡