Wherein, the first step, i.e. gradient estimation, is critical since it provides the essential direction to update variables, which have been explored by many recent works[5,30].
The first challenge is the exorbitant cost of scaling up the size of manually-labeled video datasets. The recent creation of large-scale action recognition datasets [5,15,25,26]hasundoubtedly enabled amajor leap forwardinvideo models accuracies.