This paper presents a novel approach to data attribution and error proxying for machine learning models in space missions, offering a more efficient and accurate way to assess model performance. It's a must-read for those interested in explainable AI and its applications in scientific research.
This paper introduces a new method for training data attribution in machine learning models used in space missions like ESA's Ariel. It reformulates influence in terms of prediction, computes infinitesimal prediction influence efficiently, and derives a conservative error proxy. The method is evaluated against simulated spectra and shows strong correlation with spectral errors. It also identifies influential samples and approximates harmful ones, suggesting its potential as an operational framework for scientific machine learning.
Traceable Spectral Inference via Influence Functions: Efficient Data Attribution and Error Proxies for the Ariel Mission
Interpretability is critical for machine learning models deployed in scientific space missions such as ESA's Ariel, where ground truth is unavailable during operations and physical plausibility must be assessed. While most explainable AI methods focus on feature attribution, this work investigates training data attribution through influence functions and introduces three key contributions for operational spectroscopy pipelines. First, influence is reformulated in terms of prediction rather than loss, enabling label-free deployment. Second, by leveraging the closed-form ridge solution of an Extreme Learning Machine, infinitesimal prediction influence is efficiently computed. Third, an influence-based conservative error proxy is derived by propagating training residuals through the influence sensitivities. Evaluated against simulated spectra, the proposed proxy correlates strongly with scale and shape-based spectral errors. Furthermore, influence functions enable the identification of the most influential samples and the approximation of the most harmful ones. Together, these results suggest that this approach can serve as an operational framework for scientific machine learning.