Next-Scale Transformer实现人脸多视角合成

Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

精选理由

Next-Scale Transformer模型在人脸合成方面表现出色,与现有模型结合,生成更精确的3D人脸模型,值得关注。

AI 摘要

Next-Scale Transformer模型在人脸多视角合成方面取得进展,通过高分辨率、多视角输出和跨视角一致性,实现更真实的人脸视图合成。模型在合成数据集上训练,无需2D预训练,使用通用预训练,最终生成清晰、逼真的3D人脸模型。

原文 · arXiv cs.LG

Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

Photorealistic novel view synthesis of people remains challenging at high spatial resolutions and across multiple target cameras, where preserving identity, fine appearance details, and geometric coherence is critical. We build on the next-scale autoregressive paradigm and adapt it for human-centric view synthesis by enabling higher image resolutions, multi-view outputs and stronger cross-view consistency in a single forward pass. We train on a synthetic dataset of human faces spanning diverse identities and apparel. Contrary to diffusion models, this paradigm does not need 2D pre-training and, thanks to its next-scale architecture, it benefits from lower-resolution, general-purpose pre-trainings, with the full-sized purpose-specific images being used only in the last training stages. This enables our architecture to converge with a smaller amount of purpose-specific training data, allowing us to use a smaller but more realistic training dataset. The resulting model produces sharp and realistic views, with the option to synthesize multiple novel viewpoints simultaneously for improved agreement across views. Empirically, we observe gains in perceptual fidelity and cross-view coherence on human subjects, demonstrating that next-scale autoregression is an effective backbone for scalable, multi-output human view synthesis. We also couple our pipeline with an existing transformer-based model for pixel-aligned 3D gaussian lifting from multi-view facial inputs, resulting in accurate and photorealistic 3D models of human faces.