Rethinking Dataset Distillation: 蒸馏集未必优于核心集

精选理由

想用数据集蒸馏来压缩训练集？这篇论文告诉你，现有DD方法在ImageNet上不比随机选子集好，还更贵，不如直接用核心集。

AI 摘要

这篇论文基于ImageNet-1K、ImageNet100和ImageNette三个数据集，采用三种训练协议，对七种最新数据集蒸馏（DD）方法与三种核心集选择（CS）策略进行了标准化对比。实验发现，部分DD方法甚至不如随机子集，而最先进的DD方法在大规模数据集上表现与核心集相当或更差。DD方法的构建成本显著高于CS。此外，核心集在数据分布覆盖、代表性和多样性上始终优于蒸馏集。

原文 · arXiv cs.LG

Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets?

Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples. However, DD methods are often evaluated under inconsistent evaluation protocols, ranging from standard ERM to single/multi-teacher supervision, making it difficult to isolate the effectiveness of distilled data from evaluation. Moreover, many prior methods claim that DD outperforms data pruning approaches such as coreset selection (CS), based on the assumption that restricting condensed datasets to subsets of real samples fundamentally limits their expressiveness. In this work, we critically evaluate DD methods through large-scale experiments using standardized datasets and evaluation protocols to assess their intrinsic effectiveness. We benchmark seven state-of-the-art (SOTA) DD methods on ImageNet-1K, ImageNet100, and ImageNette, using three widely adopted training protocols against three CS strategies. Our results show that while some DD methods fail to outperform even simple random subsets, the SOTA DD approaches are comparable to or worse than coresets on large-scale datasets and incur a substantially higher cost for construction. Beyond accuracy, we also evaluate the representativeness, diversity, and quality of condensed sets, and find that coresets consistently achieve better coverage of the original data distribution. These findings highlight the limited practical advantages of current DD methods and show that coresets remain competitive and are often a more computationally efficient alternative for data-centric learning.

阅读原文