搞化工的可以看看,它让LLM给符号回归当顾问,找动力学模型最多省近80%的湿实验迭代。
DASyR-LLM是一个将LLM嵌入迭代符号回归的框架,在每轮中负责批判候选模型和提出新的速率表达式。在四个虚拟案例(涵盖多相催化和生物过程)中,相比最先进的符号回归框架,找到真实模型的迭代次数减少41.7%至79.3%。超过一半的引导运行中,LLM直接提出了正确的模型结构。独立验证集上的预测性能与基线相当,所有案例R²均高于0.98。消融实验显示,符号回归组件和LLM规模共同影响性能,缩小版LLM仍保留大部分发现效率。
DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model Discovery
Kinetic model discovery is a central challenge in chemical engineering, as accurate rate expressions are essential for understanding and controlling chemical and biological processes. Symbolic regression (SR) has emerged as a powerful data-driven approach for identifying interpretable kinetic models, but usually operates without domain knowledge, often exploring physicochemically implausible models. Large language models (LLMs) offer a promising avenue for injecting domain expertise into this search. Here, we introduce an LLM-guided SR framework, embedding an LLM module within an iterative SR algorithm for automated kinetic model discovery. The LLM performs two roles at each iteration: (1) a qualitative physicochemical critique of the best SR candidates, and (2) the proposal of new candidate rate expressions guided by the SR-generated models and embedded chemical knowledge. Our framework is evaluated on four in silico case studies of increasing complexity, spanning heterogeneous catalysis and bioprocess systems. Results show the LLM-guided framework reduces iterations to identify the ground-truth model by $41.7-79.3\%$ versus a state-of-the-art SR framework, with the LLM directly proposing the correct model structure in over half of the guided runs. In practical settings, where each iteration typically requires a new wet-lab experiment, this translates into a substantial reduction in experimental effort. Predictive performance on an independent validation set is equivalent between both approaches, with $R^2>0.98$ in all case studies. Ablation studies indicate that both the SR component and the LLM scale contribute to this performance, with a reduced-size LLM largely retaining discovery efficiency. These findings demonstrate that LLMs can effectively inject domain knowledge into scientific model discovery, paving the way toward fully automated, domain-aware kinetic modelling pipelines.