模型多源确认

OpenAI 发布 MentalHealthBench 心理健康对话基准

We’re demonstrating how frontier models have continued to improve in realistic mental health convers...

精选理由

OpenAI 和 80 多位临床医生一起做了个心理健康对话评测基准,还开源了方法,做模型评测的可以拿来直接跑。

OpenAI 推出开源基准 MentalHealthBench,用于评估前沿模型在真实心理健康对话中的表现。该基准由超过 80 位心理健康临床医生参与构建。OpenAI 公开发布了整套评估方法,其他研究者可以复现评测并在其基础上继续开发。

原文 · OpenAI

We’re demonstrating how frontier models have continued to improve in realistic mental health convers...

We’re demonstrating how frontier models have continued to improve in realistic mental health conversations with MentalHealthBench. This new open benchmark was built with input from more than 80 mental health clinicians. We’re releasing it openly so other researchers can examine the methods, run their own evaluations, and build on the work. openai.com/index/introduc… 💬 129 🔄 54 ❤️ 1259 👀 100291 📊 224 ⚡