论文

Jev 1.13 医疗基准评测:准确率逊于 GPT-6 Sol 但校准更佳

Jev in Medicine: A Benchmark Evaluation. Preliminary Results

精选理由

有人拿非生成式模型 Jev 1.13 和 GPT-6 Sol 在四个医疗基准上跑了对比,便宜快十倍但复杂病例上差 20 多个百分点,概率校准数据挺有意思。

一项研究将非生成式 System One 模型 Jev 1.13 与 GPT-6 Sol 在四个医疗基准上对比,共 8,469 次请求全部返回有效答案。Jev 在 PubMedQA 上与 GPT-6 Sol 持平(78.4% vs 78.2%),在 MetaMedQA 上为 74.8% vs 82.7%,在 DiagnosisArena-MCQ(59.8% vs 82.4%)和 NEJM 病例挑战(61.8% vs 82.4%)上明显落后。Jev 的概率校准最好,期望校准误差为 0.063(GPT-6 Sol 为 0.146),高置信答案(概率≥0.9,占 52.9%)准确率达 93.4%。其延迟中位数为 0.27-0.31 秒,2,823 道题总成本仅 0.08 美元,作者强调临床使用前需针对具体任务验证。

原文 · arXiv cs.LG

Jev in Medicine: A Benchmark Evaluation. Preliminary Results

Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.