研究实测 JEV 与 LLM 在七项政治学复现任务中的准确率与成本
JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications
想用模型标注文本的研究者可以看看:JEV 号称便宜快,实测发现便宜不成立,校准也只赢了一半。
arXiv 论文对比了 TypeSafe 推出的 System One 类模型 JEV 与 GPT-6 Luna、Qwen3.8-27B 在七个政治学文本复现任务上的表现。结果显示 JEV 的准确率接近两个 LLM,但按 OpenAI 批处理价格计算并无成本优势。单次提问时 JEV 的概率校准优于 GPT-6 Luna 的 token 概率,但不稳定地优于 Qwen3.8-27B。论文结论是除非研究者特别需要速度,JEV 的明显优势只剩解析选择概率更方便。
JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications
Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe markets as "System One" models, instead returns decisions and probability distributions across a user-supplied fixed answer set. A commercial model, JEV, is advertised as having a dramatic cost and speed advantage over traditional LLMs along with better calibrated decisions. As such, it might be useful for social scientists looking to quickly and cost-effectively annotate or scale large corpora of text and have a reliable indicator of a classifier's uncertainty. Yet, the accuracy of these claims and the broader model accuracy in social science text-based tasks are not yet established. In this paper, we do just that and hope to establish the suitability of JEV for social science tasks. We compare JEV with LLMs and human coders from published research, and with a current mid-tier commercial LLM (GPT-6 Luna) and an open-weight alternative (Qwen3.8-27B). We find that JEV matches, or comes close to, the capabilities of both LLMs in a variety of tasks. However, we find no cost advantage over GPT-6 Luna at OpenAI's batch prices. Further, we find that, when each question is asked once, JEV's probabilities are better calibrated than GPT-6 Luna's token probabilities, but not consistently better than Qwen3.8-27B's. We conclude that unless researchers have a need for speed, JEV's only obvious advantage is ease of parsing the underlying choice probabilities.