共识伪标签学习新方法:通过类内结构调整构建差异化决策源
Constructing Structured Decision Sources for Consensus-Based Pseudo-Label Learning
这篇论文说伪标签共识想有效,得让模型看不同的证据。他们在 Cora、CiteSeer、PubMed 上把伪标签精度最多提了 4.39 个百分点,做图学习可以看看。
arXiv 论文提出在共识伪标签学习中通过控制类内结构变化来构建决策源,而非简单堆叠模型。方法基于共享图表示,改变中心粒度与邻域混合方式,并用节点对共分配选出互补子集。在 Cora、CiteSeer、PubMed 三个公开数据集上,5 个随机种子下,固定预算训练的伪标签精度比三个常规初始化 GCN 源提升 1.19 到 4.39 个百分点;在匹配结构过滤下,三源共识在全部 45 个数据集槽位对比中精度均高于任一单源。下游准确率有竞争力但并非在每个数据集上都领先。
Constructing Structured Decision Sources for Consensus-Based Pseudo-Label Learning
Consensus can make pseudo-label learning more reliable, but only when its predictors contribute genuinely different evidence. Multiple models that repeat the same boundary provide additional votes without additional information. We address this problem by con structing decision sources through controlled changes to within-class structure. Starting from a shared graph representation, we vary center granularity and neighborhood mixing, reproduce each resulting source to test its stability, and select a complementary subset using node pair coassignment. Unanimous predictions from the selected sources are then ranked for student training. On the public fixed splits of Cora, CiteSeer, and PubMed, evaluated with five random seeds, the constructed sources improve fixed-budget training pseudo-label precision by 1.19 to 4.39 percentage points over three conventionally initialized GCN sources. Under matched structural filters, three-source consensus is more precise than each constituent source in all 45 dataset slot eed comparisons. The gains are strongest in pseudo-label quality: downstream accuracy remains competitive but does not lead on every dataset. These results identify source construction rather than model count alone as an important design problem for consensus-based pseudo-label learning.