论文

SPEAR:谱解耦 MoE 神经算子,面向大规模 PDE 预训练

SPEAR: A Spectral-Disentangled MoE Neural Operator with Knowledge-Guided Expert Aggregation for Large-Scale PDE Pretraining

精选理由

一篇解决 PDE 基础模型知识干扰和 MoE 专家冗余的论文,专家数量砍一半精度不掉,做科学计算的可以看看。

SPEAR 是一个用于大规模 PDE 预训练的谱解耦 MoE 神经算子,将潜在特征拆分为低频与高频分量,低频部分共享建模可迁移动力学,高频部分专门学习 PDE 特有模式。针对 MoE 专家冗余问题,它提出知识引导的专家聚合策略,基于数据集知识表示与路由偏好度量专家相似度并合并相似专家。实验覆盖 12 个 PDE 数据集及多个下游基准,在预训练、微调和迁移学习上表现优于现有方法。该聚合策略可将专家数量减少 50%,同时保持或提升预测精度。

原文 · arXiv cs.AI

SPEAR: A Spectral-Disentangled MoE Neural Operator with Knowledge-Guided Expert Aggregation for Large-Scale PDE Pretraining

Large-scale pre-training has improved the generalization of neural operators across diverse PDEs. However, existing PDE foundation models still struggle with heterogeneous dynamics, where shared representations may cause knowledge interference, while mixture-of-experts (MoE) architectures suffer from increasing expert redundancy. We propose SPEAR, a spectral-disentangled MoE neural operator with knowledge-guided expert aggregation for large-scale PDE pre-training. SPEAR decouples latent features into low- and high-frequency components, enabling shared modeling of transferable dynamics and specialized learning of PDE-specific patterns. To address expert redundancy, we design a knowledge-guided expert aggregation strategy that measures expert similarity from dataset-specific learned knowledge and routing preferences, enabling the identification and consolidation of similar experts. Experiments on twelve PDE datasets and multiple downstream benchmarks demonstrate superior performance in pre-training, fine-tuning, and transfer learning. Furthermore, our aggregation strategy reduces the number of experts by 50\% while maintaining or improving prediction accuracy, achieving a balance between model efficiency and generalization for PDE foundation models.