DiagLoop只用合成数据就把8B诊断模型练到超过专有模型,工业系统+11.6分,疾病分类+5.5分,可以看看它的反事实数据飞轮。
DiagLoop将物理关系或临床指南转化为训练数据,通过反事实世界生成与独立检查器构成数据飞轮。DiagLoop用阶段局部强化学习只更新模型生成的延续部分,并用同一标准控制准入和奖励。仅用合成场景训练,8B模型在八个工业系统上严格路径正确性提升11.6点,在十个疾病类别上提升5.5点。相对对照提升分别为3.9和2.3点,且超过评测的专有参考模型。
DiagLoop: A Counterfactual Data Flywheel with Stage-Localized Reinforcement for Diagnostic LLMs
Causal diagnostic models must explain how conclusions follow from evidence because diagnoses guide repairs and treatments. Yet serious cases are scarce, records rarely contain reasoning paths, and data transfer poorly across configurations, complicating local deployment. We present DiagLoop, a counterfactual data flywheel that converts codified physical relations or clinical guidelines, authored once per mechanism family, into training supervision beyond recorded cases. A training-only teacher proposes counterfactual worlds by varying causes, contexts, and observations, while an independent hybrid checker admits only valid worlds. The student reasons through symptom abstraction, causal-chain construction, and root-cause attribution. Stage-specific criteria identify its earliest failure. For nonterminal failures, a bounded repair probes downstream competence, and the resulting weakness profile guides subsequent data generation. Stage-localized reinforcement learning updates only the model-generated continuation, while replay and preservation reduce forgetting. The same criteria govern admission, attribution, reward, and regeneration through checks separate from the proposer. Using only synthesized scenarios and no case-level expert reasoning annotations, the resulting 8B model improves strict path correctness over the strongest conventional baseline. Gains are 11.6 points across eight industrial systems and 5.5 points across ten disease categories. Gains over a deranged-routing control are 3.9 and 2.3 points, respectively. The model also exceeds the evaluated proprietary references in both domains, even when they receive few-shot examples or the specification in context.