这篇arXiv论文搞了个新架构,让AI不碰物理世界也能跨领域找数学等价关系,还自带DAB-30基准能检验,值得搞科学发现的朋友看看。
论文论证在线具身并非所有科学溯因的必要条件,提出身份溯因可通过表征接地实现。作者设计Abduction Loop架构,包含表示生成、主题提取、约定空间规范化、跨域检索等环节,并以放弃为默认输出。文中以多模态模型从引力记忆输运模型图生成并验证了与弱引力透镜宇宙学的球形Kaiser-Squires质量映射等价的假设作为示例。作者提出可证伪的评估基准DAB-30,包含30个任务,用于检验该机制的能力边界。
Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation
Can scientific abduction occur without continuous sensorimotor embodiment? Recent arguments in AI and philosophy of science hold that genuine hypothesis generation requires an agent continuously coupled to the physical world. We defend a narrower claim: online embodiment is not necessary for every abductive scientific act. Our focus is identity abduction: the inference that two independently developed structures are one object under an explicit correspondence, reached through representational grounding rather than bodily interaction. An agent may acquire new inferential affordances not through physical interaction but through transformations into representations that expose latent invariants. Scientific diagrams are a practical substrate because they embody independently evolved conventions that partially canonicalize symmetry, topology, and operator structure across disciplines - a property we develop as convention space, which answers a hard retrieval problem: finding mathematically related work when two fields share no discriminating vocabulary. We operationalize the mechanism as an architecture, the Abduction Loop: representation generation, motif extraction, convention-space canonicalization, cross-domain retrieval, identity-hypothesis generation, and adversarial verification, with abstention as the designed default. A documented episode, in which a multimodal model given a figure of a gravitational-memory transport model generated and then verified the hypothesis that its central differential complex is equivalent to the spherical Kaiser-Squires mass-mapping complex of weak-lensing cosmology, serves as a motivating possibility witness from which the architecture is abstracted, not as evidence of general capability. We close with a falsifiable evaluation program, the DAB-30 benchmark. The contribution is a mechanistic proposal, an architecture, and a test program.