MEND:世界模型中的潜在幻觉检测与修正
MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models
MEND方法能检测并修正世界模型中的幻觉,AUROC达0.80,定位准确率0.87,提升预测准确性。
研究人员提出MEND方法,用于检测、定位和修正世界模型中的潜在幻觉现象。该方法在两个导航环境中实现了0.80的AUROC值,无需使用动作即可检测幻觉。MEND通过单一条件评分网络实现,能够将错误定位到特定图像块,并定义推理时的修正方向。研究还发现部分错误与数据流形相切,因此重点强调了检测和定位的成果。
MEND: Label-Free Detection, Localisation, and Correction of Latent Hallucination in World Models
World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the phenomenon of hallucination in latent World Models: given a state and an action, the predicted next latent can decode to a scene that never occurs. Because the prediction is statistically ordinary and is fed back autoregressively by the model, the error is both silent and compounding. We study whether such latent hallucination can be detected, localised, and corrected at inference time, on a frozen self-supervised world model in the absence of ground-truth error labels. We introduce Masked Empirical-Bayes Neural Denoising (MEND), a single conditional score network trained by denoising score matching on real transitions, whose score field serves three roles: its magnitude detects hallucination, its per-token field localises it to specific image patches, and it defines an inference-time correction direction. On two navigation environments MEND detects hallucination with an AUROC of up to 0.80 without using actions, exceeding a single-Gaussian density baseline while also localising the error (per-token AUPRC up to 0.87) and correcting it, all from one score field. Our correction reliably reduces single-step latent error and improves predictions. We identify that a part of the error is tangent to the data manifold, hence, we focus on detection and localisation while highlighting promises of the correction.