模型多源确认83°

OpenAI 发布 MentalHealthBench,GPT-6 Astra 得分 57.3

精选理由

OpenAI 拉了 80 多位心理医生做了个心理健康对话基准,GPT-6 Astra 考了 57.3 分,GPT-4o 只有 32.1,想看模型心理对话靠谱程度可以关注这套评测。

OpenAI 联合来自 22 个国家的 80 多名执业心理医生和精神科医生,推出了心理健康对话评测基准 MentalHealthBench。该基准覆盖 19 种语言和近 20 个细分专业方向,专注日常及模糊情境的心理对话,而非以往的紧急危机场景。每段合成对话都配有专家定制评分标准,至少 3 名专家审核,需 2 人同意且第 3 人不反对才保留。在新基准上,GPT-6 Astra 得分 57.3,GPT-4o 得分 32.1。评分由 GPT-5.6 Sol 对照人类撰写的标准进行,因此质量仍部分依赖 LLM 评委。

原文 · rohanpaul_ai

OpenAI's new MentalHealthBench puts GPT-6 Astra at 57.3, versus GPT-4o's 32.1, on realistic mental health conversations.

Most mental health AI evaluations have centered on emergencies and broad safety criteria, leaving everyday and ambiguous conversations much less measured.

So OpenAI co-created this benchmark with more than 80 licensed psychologists and psychiatrists from 22 countries, spanning 19 languages and nearly 20 subspecialties.

Each synthetic conversation gets a custom expert rubric covering behaviors such as seeking context, preserving user agency, safety, and appropriate guidance.

At least 3 experts reviewed each case, and a criterion survived only when 2 agreed and a 3rd did not contradict it.

GPT-5.6 Sol then grades model answers against those human-written criteria, so score quality still partly depends on an LLM judge.