想从单目视频中重建4D动态场景?SM4RT用刚体运动基分解运动,比独立点流更物理合理,结果更准确。
SM4RT提出了一种结构化运动4D重建Transformer,将场景运动分解为紧凑的运动基,每个基表示为SE(3)中的6D扭转时序。不同于现有方法将运动视为独立点位移,SM4RT利用刚体运动学,通过稀疏、时间共享的逐像素分配权重恢复稠密场景运动。在单目RGB视频输入下,SM4RT联合推断3D几何、世界坐标运动和运动学结构。实验表明,SM4RT在4D重建任务上取得了强劲的运动重建性能,同时保留了场景运动的几何结构。
SM4RT: Learning Structured Motion Geometry for 4D Reconstruction
Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding remains a fundamental challenge. Most existing motion perception methods (e.g., sparse tracking, dense point-wise flow) treat motion as independent point-wise displacements, ignoring the structured nature of physical motion. However, real-world objects usually obey rigid-body kinematics, and points thus usually move collectively, not in isolation. Motion itself possesses geometric structure: physical objects undergo a set of rigid-body transformations governed by SE(3), rather than unstructured point-wise displacements. Building on this insight, we propose SM4RT, a Structured Motion 4D Reconstruction Transformer for end-to-end 3D reconstruction and structured motion perception. SM4RT introduces Structure-of-Motion to represent scene dynamics, where scene motion is decomposed into a compact set of motion bases, each represented as a temporal sequence of 6D twists in SE(3). Dense scene motion is then recovered by sparse, time-shared per-pixel assignment weights over these bases, ensuring points on the same object share a common rigid-body motion trajectory. SM4RT introduces a parallel motion geometry encoder and decoder that jointly infer 3D geometry, world-coordinate motion, and scene kinematic structure in a single forward pass from monocular RGB video. SM4RT achieves strong motion reconstruction performance while preserving the geometric structure of scene motion.