Papers
arxiv:2609.33253

VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis

Published on Sep 27
· Submitted by
chenkangjie1123
on Sep 29
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide strong generative priors yet rely on implicit source-to-query correspondence. VGGT-Diff bridges these regimes by routing visual geometry latents from VGGT-Ω into a pretrained video diffusion model. Each visual token is associated with a 3D point and confidence, then transformed into query-aligned latent conditions through a confidence-aware Visual Geometry Router (VGR) that preserves front and back surface evidence. These conditions guide joint target-view denoising, while Point-Track Residual Consistency (PTRC) regularizes predicted-clean residuals along reliable 3D tracks, improving multi-view stability. We further introduce robust geometry conditioning, combining training-time regularization with inference-time guidance for improved robustness. Experiments show competitive or state-of-the-art performance across interpolation and extrapolation under different viewpoint difficulties. Our code is available at https://github.com/chenkangjie1123/VGGT-Diff.

Community

Paper author Paper submitter
•
edited 12 days ago

VGGT-Diff_Figure1_teaser_600dpi

VGGT-Diff is a geometry-routed multi-view diffusion model for sparse-view novel view synthesis from six input images. Its first key innovation, the confidence-aware Visual Geometry Router (VGR), transforms VGGT-Ω features into query-aligned geometric conditions while preserving both front- and back-surface evidence. Its second innovation, Point-Track Residual Consistency (PTRC), regularizes denoising residuals along reliable 3D tracks to improve cross-view geometric consistency.

Despite being trained on only 1K scenes with a relatively limited training budget, VGGT-Diff already delivers competitive or state-of-the-art novel-view synthesis performance. It achieves state-of-the-art PSNR / LPIPS in different viewpoint-difficulty settings, with particularly strong results on challenging mid- and far-range interpolation and extrapolation views. Further scaling in both training data and optimization steps could continue improving visual quality and cross-view consistency.

VGGT-Diff supports both pose-aware inference with known camera parameters and pose-free inference directly from six RGB images. The current effective 21K checkpoint—initialized from a 20K half-resolution checkpoint and adapted with only 1K additional full-resolution training steps—already generates high-quality 480p, 80-frame continuous camera-trajectory videos. The model is still being actively trained, and we plan to release multiple fully trained checkpoints to the community later.

Good work!

·
Paper author

Thks!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33253
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 3

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33253 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33253 in a Space README.md to link it from this page.

Collections including this paper 1