如果你的时间序列模型总是泛化不好,试试用RMISC这个真实世界大数据集预训练,1420亿个时间点,效果比合成数据强很多。
RMISC语料库包含约200个数据集、1420亿个时间点,覆盖多领域真实世界多变量时间序列。研究者用RMISC预训练四种先进时间序列基础模型(TSFMs),并与合成数据预训练模型对比。实验表明,真实世界多变量数据显著提升TSFMs的零样本泛化能力,在分布内和分布外基准上均表现更优。
RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models
Recent years have witnessed the emergence of multivariate modeling using time series foundation models (TSFMs), which achieve advanced zero-shot generalization. Modern multivariate TSFMs are predominantly pretrained on multivariate synthetic data, which is easier to scale but may fail to capture the complex temporal dynamics and cross-variable relationships present in real-world time series. This raises a key question: Whether and to what extent the leading TSFMs trained with the real-world corpus perform better than those trained with synthetic data? To answer this, we establish the RMISC corpus, a considerably large-scale, high-quality, openly accessible, real-world, and multivariate time series archive that contains around 200 datasets and 142 billion time points across diverse domains. Furthermore, we pretrain four advanced TSFMs on univariate, synthetic multivariate, and real-world multivariate data and evaluate their zero-shot generalization capabilities on standard in-distribution and out-of-distribution benchmarks. Experimental results show that incorporating real-world multivariate data predominantly improves the generalization performance for both univariate and multivariate TSFMs. These results provide a deeper understanding of how real-world multivariate data contributes to the development of stronger TSFMs.