想测语言模型能不能真做化学实验?onepot-Bench 0 用私有数据考反应预测和拒毒能力,比公开题更贴近湿实验室。
onepot-Bench 0 是面向湿实验室合成化学能力的语言模型评测基准,包含三项子评估。ChemAbacus 考查无需工具的化学信息学素养与数值推理;SynthRefusal 评测模型对苯二氮卓类、受控药物及设计药物等目标的拒绝与安全行为;SynthBench 使用实验室私有实验数据评测反应结果预测和催化剂选择。该基准旨在补充现有评测中缺失的物理实验室可靠决策能力。
onepot-Bench 0: towards lab-aware in silico chemistry benchmarks
Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc analysis. However, precisely measuring their abilities is difficult, as scientific capabilities require a mixture of both problem-solving skills and domain-specific intuition. Existing evaluations rarely measure the capabilities required to make reliable decisions in a physical laboratory and often rely on public data that may have appeared in model training corpora. We introduce onepot-Bench 0, a proprietary benchmark suite for evaluating language models on synthetic chemistry capabilities relevant to wet-lab execution. onepot-Bench 0 comprises three complementary evaluations: ChemAbacus measures tool-free cheminformatics literacy and numerical reasoning; SynthRefusal characterizes safety and refusal behavior across a variety of benign, controlled, and designer-drug targets; and SynthBench evaluates reaction-outcome prediction and catalyst selection using private experimental data generated in our laboratory. Together, these evaluations probe basic competency, reliability, and deeper knowledge, all skills which are required for reliable performance in the lab.