机器学习模型验证:从留出法到嵌套交叉验证
Model validation in machine learning: A scenario-based guide from hold-out splits to nested group cross-validation in biomedical and applied research
机器学习研究者必读,教你如何避免数据泄露陷阱,选择正确的验证方法。
这篇教程详细介绍了机器学习模型验证的多种方法,包括留出验证、k折交叉验证和嵌套交叉验证等。研究通过8个受控场景比较了有缺陷和泄漏安全的设计方案,涉及EEG时段、配对眼部OCT图像和临床测量数据。作者提供了MATLAB和scikit-learn的可复现模板,以及决策树和报告检查清单。
Model validation in machine learning: A scenario-based guide from hold-out splits to nested group cross-validation in biomedical and applied research
Model validation estimates the performance of a complete learning procedure on new data. However, an invalid split can produce an optimistic and stable result. This tutorial reviews hold-out validation, train/validation/test designs, repeated random subsampling, k-fold and repeated stratified cross-validation, leave-one-out and leave-p-out schemes, group-aware validation, and nested group cross-validation. General machine-learning principles are linked to EEG epochs, paired-eye OCT images, repeated clinical measurements, and multicenter data. Eight controlled scenarios compare flawed and leakage-safe designs: seven use locked confusion matrices with auditable metrics, and one uses a reproducible repeated-study simulation. The scenarios cover global feature selection, normalization leakage, dependent records, center mixing, repeated test-set use, and estimator instability. Bias, variance, metric aggregation, uncertainty, and computational cost are also examined. A data-size matrix, a decision tree, and reporting checklists are provided. Reproducible MATLAB templates and scikit-learn counterparts are included. The results show that no validation method is universally best. The independent unit must match the intended deployment target. Every data-dependent operation must also exclude the observations used for performance estimation.