Novel View Synthesis with Diffusion Models
Watson, Daniel, Chan, William, Martin-Brualla, Ricardo, Ho, Jonathan, Tagliasacchi, Andrea, Norouzi, Mohammad
–arXiv.org Artificial Intelligence
We present 3DiM, a diffusion model for 3D novel view synthesis, which is able to translate a single input view into consistent and sharp completions across many views. The core component of 3DiM is a pose-conditional image-to-image diffusion model, which takes a source view and its pose as inputs, and generates a novel view for a target pose as output. The output views are generated autoregressively, and during the generation of each novel view, one selects a random conditioning view from the set of available views at each denoising step. We demonstrate that stochastic conditioning significantly improves the 3D consistency of a naïve sampler for an image-to-image diffusion model, which involves conditioning on a single fixed view. We compare 3DiM to prior work on the SRN ShapeNet dataset, demonstrating that 3DiM's generated completions from a single view achieve much higher fidelity, while being approximately 3D consistent. We also introduce a new evaluation methodology, 3D consistency scoring, to measure the 3D consistency of a generated object by training a neural field on the model's output views. Given a single input image on the left, 3DiM performs novel view synthesis and generates the four views on the right. We trained a single 471M parameter 3DiM on all of ShapeNet (without classconditioning) and sample frames with 256 steps (512 score function evaluations with classifier-free guidance). See the Supplementary Website for video outputs.
arXiv.org Artificial Intelligence
Oct-6-2022
- Country:
- North America > Canada
- Asia > Japan
- Honshū > Chūbu
- Nagano Prefecture > Nagano (0.04)
- Ishikawa Prefecture > Kanazawa (0.04)
- Honshū > Chūbu
- Genre:
- Research Report (1.00)
- Technology: