论文

自监督学习用于极光光谱:1D ViT 掩码预训练超越专家特征与监督基线

Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra

精选理由

用 22 万条极光光谱做掩码预训练,不用标签就超过专家手工特征和监督基线,还对比了 SpectraFM 和 SpecFormer 的可迁移性,做科学数据自监督的可以看看。

ASIS 极光光谱仪记录的数十万条光谱中,只有几百条有专家标注。研究用掩码自编码器在 223,000 条无标注光谱上预训练 1D Vision Transformer,其表征可恢复诊断沉降粒子的发射线强度比(R² 0.91,未训练对照为 0.77),单个线性探针的分类效果等同 13 个专家手工特征。微调后在原监督分类器自己的基准上达到 macro-AP 88.5(对手 77.8),mAP 0.870,仅用 10% 标签即比从头训练高 0.159。迁移测试显示光谱窗口决定可迁移性:红外训练的 SpectraFM 低于未训练对照,光学训练的 SpecFormer 接近但未达到域内预训练水平。

原文 · arXiv cs.LG

Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra

Auroral spectrographs such as the Auroral Spectrograph In Skibotn (ASIS) record hundreds of thousands of emission spectra, but only a few hundred can be labelled by an expert. To exploit the rest, we pretrain a 1D Vision Transformer with a masked autoencoder on 223,000 unlabelled spectra. Without labels, its representation recovers the emission-line intensity ratios that physicists use to diagnose the precipitating particles (R^2 0.91 vs. 0.77 for an untrained control) and, under one linear probe, classifies as well as 13 features designed by experts. Fine-tuned, the model outperforms the previous supervised auroral classifier on its own benchmark (macro-AP 88.5 vs. 77.8), reaches 0.870 mAP, and exceeds the same architecture trained from scratch by +0.159 with 10% of the labels; attribution shows that it uses both N2+ bands. Could an existing pretrained model replace it? Two astronomical spectral foundation models and a time-series model transfer according to their spectral window: SpectraFM, trained in the infrared, falls below the untrained control, whereas SpecFormer, trained in the optical, approaches in-domain pretraining without reaching it.