用验证器帮语言模型修量子场论对偶声明,效果不错,deepseek-chat提升8个百分点,qwen-plus提升7个百分点,值得关注。
DualityCert是一个符号验证器,用于评估四维N=1 quiver规范理论中候选Seiberg对偶声明的't Hooft异常匹配、超势R-荷一致性、中心荷匹配和有界手性环代理。在145个破损声明的预注册基准上,验证器门控重试将deepseek-chat的最终修复成功率提高了8.3个百分点(pp),将qwen-plus提高了7.1 pp(Holm调整p<0.002)。在11次尝试的等预算下,deepseek-chat上停-先策略组合比独立验证器过滤重采样低10.3 pp,而qwen-plus上则高14.7 pp,两种验证器利用策略在两个确认模型上顺序反转。类别级验证器反馈在qwen-plus上比无内容重试高8.7 pp,可解释义务标识比结构相同的掩码反馈高6.4 pp,而deepseek-chat上未检测到任一效果。
DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory
We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a bounded chiral-ring proxy. A claim that passes receives a consistency certificate, which states that no tested inconsistency was found, not that the duality is proven. We use the verifier as a repair environment for language-model agents, which receive a deliberately broken claim and must edit it until it certifies. On a preregistered benchmark of 145 broken claims, with the analysis fixed before the first confirmatory model call, verifier-gated retry improves final repair success over a single attempt by +8.3 percentage points (pp) on deepseek-chat and +7.1 pp on qwen-plus (Holm-adjusted p<0.002). Under an equal budget of eleven attempts, the stop-first strategy portfolio underperforms independent verifier-filtered resampling by 10.3 percentage points on deepseek-chat but outperforms it by 14.7 points on qwen-plus, reversing the ordering of the two tested verifier-exploitation policies across the two confirmatory models. On qwen-plus, category-level verifier feedback is worth +8.7 pp over content-free retry, and interpretable obligation identities alone are worth +6.4 pp over structurally identical masked feedback. Neither effect is detected on deepseek-chat. Separately, a preregistered MiniMax-M2.5 extension again finds an iteration gain and independent verifier-filtered resampling outperforming the strategy portfolio. Which policy is better thus differs between the two models, while every winning policy uses the same cheap certificate. The verifier, benchmark, protocol, and all per-attempt records are released.