神经网络复杂度能否提升健康虚假信息检测?一项跨语料基准研究
Does Neural Complexity Improve Health Misinformation Detection? A Leakage-Controlled Cross-Corpus Benchmark
这篇论文拿 5 种神经网络和线性 SVM 掰手腕,结果在 CONSTRAINT 上 SVM 反而赢了集成模型,做分类任务选型前可以看看。
该研究在 COVID19-FNIR 和 CONSTRAINT 两个语料库上,对 1D-CNN、LSTM、BiLSTM、CNN-LSTM、CNN-BiLSTM 五种神经网络、一个软投票集成模型和三个经典机器学习基线做了泄漏控制的对比基准。在 COVID19-FNIR 上,深度集成模型 macro-F1 达 0.9963,各神经网络在 0.9945 到 0.9957 之间。但在 CONSTRAINT 上,线性 SVM 以 macro-F1 0.9574 超过集成模型的 0.9272。架构排名随语料库变化,简单稀疏线性模型依然很有竞争力,说明基准构建方式对结果的影响可能大于架构选择。
Does Neural Complexity Improve Health Misinformation Detection? A Leakage-Controlled Cross-Corpus Benchmark
Increasing architectural complexity is often assumed to improve health misinformation detection, yet reported gains are difficult to interpret when studies use different corpora, preprocessing pipelines, data splits, and leakage controls. This study provides a controlled cross-corpus benchmark of five compact neural architectures (1D-CNN, LSTM, BiLSTM, CNN-LSTM, and CNN-BiLSTM), a soft-voting neural ensemble, and three classical machine-learning baselines using COVID19-FNIR and CONSTRAINT. Exact-text duplicate controls were applied before modelling; all neural systems used a common preprocessing and optimisation protocol, and neural results were repeated across three random seeds. On COVID19-FNIR, the deep ensemble achieved a mean macro-F1 of 0.9963 and ROC-AUC of 0.9994, while individual neural models ranged from 0.9945 to 0.9957 macro-F1. On CONSTRAINT, the ensemble achieved macro-F1 of 0.9272 and ROC-AUC of 0.9811, whereas a linear SVM achieved macro-F1 of 0.9574 and ROC-AUC of 0.9931. Architecture rankings changed across corpora, and simple sparse linear models remained highly competitive. The findings show that model complexity does not provide a stable performance advantage and that benchmark construction can dominate architecture choice. The study contributes a reproducible, leakage-controlled basis for evidence-driven model selection in health misinformation classification