OmniScientist:全模态全学科AI科学家

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

精选理由

OmniScientist直接读图像、视频、3D原始数据,36案例全跑通,均分6.3,对比实验85%胜出。

AI 摘要

OmniScientist是端到端的多模态AI科学家系统,可直接处理图像、信号、音频、视频、3D结构等原始证据。系统由感知层和三个自主智能体(构思、实验、写作)组成,通过代码检查保障新颖性与统计有效性。在覆盖5个学科家族、4类证据的36个真实数据案例中,系统全部完成从原始数据到编译手稿,平均论文得分为6.3。与仅接收预计算标量特征的盲变体相比,直接感知在7个评估维度全部占优,并在85%的成对比较中胜出。

原文 · arXiv cs.AI

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.