这篇论文给了你从数据识别ODE的理论底线,告诉你最少需要多少观测才能唯一确定方程,做科学机器学习的必读。
本文引入解集上的Hausdorff距离作为比较微分方程的自然度量,该度量捕捉两个方程在所有初始条件下的最坏情况分离,从而编码了识别问题的极小极大结构。作者建立了线性和非线性(Lipschitz/Hölder连续向量场)ODE的可识别性边界,明确了何时能从解数据中区分两个不同方程。利用该度量,推导了相关ODE类的度量熵估计,并量化了可靠恢复控制方程所需解观测的样本复杂度界限。
Recovering Governing Equations from Solution Data: Identifiability Bounds for Linear and Nonlinear ODEs
Learning governing equations from observed solution data is a fundamental challenge in scientific machine learning \cite{bruntonDiscoveringGoverningEquations2016,kovachkiNeuralOperatorLearning2023,longPDENetLearningPDEs2018,rudyDatadrivenDiscoveryPartial2017,raonicConvolutionalNeuralOperators2023}, yet the theoretical conditions under which a ground-truth ODE can be uniquely and stably identified from multiple solution observations remain largely undeveloped, and no quantitative analysis of the sample complexity of such learning tasks exists in the literature. To address this gap, we introduce the Hausdorff distance on solution sets as the natural metric for comparing differential equations, since it captures the worst-case separation between two equations over all admissible initial conditions and thus encodes the minimax structure of the identification problem. We establish identifiability bounds for governing ODEs across a wide class of structure equations--ranging from linear ODEs to nonlinear classes with Lipschitz (Hölder)-continuous vector fields--characterizing precisely when two distinct equations can be distinguished from solution data. Using this metric, we derive metric entropy estimates for the relevant ODE classes and analyze sample complexity bounds, quantifying how many solution observations are needed to reliably recover the governing equation.