MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and Interpolation

Feb-5-2026, 17:00:30 GMT–Neural Information Processing Systems

Video prediction is a challenging task. The quality of video frames from current state-of-the-art (SOTA) generative models tends to be poor and generalization beyond the training data is difficult. Furthermore, existing prediction frameworks are typically not capable of simultaneously handling other video-related tasks such as unconditional generation or interpolation. In this work, we devise a general-purpose framework called Masked Conditional Video Diffusion (MCVD) for all of these video synthesis tasks using a probabilistic conditional score-based denoising diffusion model, conditioned on past and/or future frames. We train the model in a manner where we randomly and independently mask all the past frames or all the future frames.

artificial intelligence, machine learning, masked conditional video diffusion, (7 more...)

Neural Information Processing Systems

Feb-5-2026, 17:00:30 GMT

Conferences Web Page

Add feedback

Technology:
- Information Technology > Artificial Intelligence > Machine Learning (0.96)