做 3D 内容生成或多视图编辑的开发者,终于有了能处理大幅几何变形的工具——GeM-NR 无需训练即可与主流编辑器配合,建议试试看能否解决你场景中的非刚性编辑痛点。
现有多视图图像编辑方法大多局限于刚性或外观编辑,无法处理改变场景几何的非刚性编辑。GeM-NR 提出了一种无需训练的快速方法,通过深度图对齐、视角投影和条件细化,实现多视图一致的几何与外观编辑。该方法兼容 FLUX、Qwen、BrushNet 等主流编辑器,支持从两视图扩展到多视图,显著提升了编辑质量和几何光度一致性。实验表明,GeM-NR 在非刚性编辑任务上达到当前最优水平,甚至能生成编辑后的 3D 表示。
GeM-NR: Geometry-Aware Multi-View Editing for Nonrigid Scene Changes
Recent developments in multi-view image editing with generative models have brought us a step closer toward general 3D content generation and customization. Most existing works focus on rigid or appearance-only edits by utilizing the geometry of the unedited scene. This naturally limits these methods to edits that preserve the underlying scene structure. Other approaches are trained for specific image editing tasks, such as object removal and addition. Despite this progress, general nonrigid edits, i.e., edits that substantially change the scene geometry, remain challenging for existing methods. We propose GeM-NR, a fast and flexible training-free approach for general multi-view consistent image editing, including edits that drastically change the geometry and appearance of the scene. Given an anchor image edited with a chosen backbone editor (such as FLUX, Qwen, BrushNet) and a query unedited image, GeM-NR edits the query image consistently with the anchor edit. The method incorporates multiple stages: (i) depth map estimation, where we propose a strategy to maximize the alignment between the 3D point clouds of the edited and unedited scenes, (ii) projection onto a query viewpoint, and (iii) refinement of the obtained image conditioned on the unedited query. The conditioning-based formulation scales well from two to many views of an object. We demonstrate the ability of our method to handle edits with significant changes in geometry and appearance, something that existing methods struggle with. We perform an extensive evaluation showing that our method improves consistency for a wide variety of edit tasks, including generating 3D representations of the edited scene. Both quantitative and qualitative results indicate the state-of-the-art performance of our method in terms of edit quality as well as geometric and photometric consistency across multiple views.