ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling

Neural Information Processing Systems 

We propose ID-to-3D, a method to generate identity-and text-guided 3D human heads with disentangled expressions, starting from even a single casually captured'in-the-wild' image of a subject. The foundation of our approach is anchored in compositionality, alongside the use of task-specific 2D diffusion models as priors for optimization. First, we extend a foundational model with a lightweight expression-aware and ID-aware architecture, and create 2D priors for geometric and texture generation, via fine-tuning only 0.2% of its available training parameters.