Long-Term Prediction of Natural Video Sequences with Robust Video Predictors
–arXiv.org Artificial Intelligence
While looking at a static image alone it is difficult to discern (without direct supervision) between background and foreground objects, or discern individual objects at all. Where does one object start and another end? Only by observing the objects interacting with each other over time can we fully learn their properties. Prediction of video sequences using only video frames is especially difficult as images are a discrete 2D projection of the real world. Many difficulties in frame-to-frame video prediction come from the fact that our system must perform several sub-tasks sequentially in order to create an output that not only looks realistic, but also matches real world data. We argue that these two requirements, predicting a realistic frame and predicting a frame that matches a target frame from the data-set, conflict with each other during training causing poor performance. In order to create plausible predictions in dynamic natural video sequences, we loosen the constraint of matching the target frames exactly. By doing so, we acknowledged that there are some things in the next frame that are not possible to predict at all, let alone with any accuracy.
arXiv.org Artificial Intelligence
Aug-21-2023
- Country:
- Europe > Italy > Calabria > Catanzaro Province > Catanzaro (0.04)
- Genre:
- Research Report (0.64)
- Technology: