Deep Learning
Appendix for Integrating Momentum into Recurrent Neural Networks
Section 3.1, we flatten and process the image as a sequence of the length of 784 pixel-by-pixel. The baseline LSTM models consist of one LSTM cell with 128 and 256 hidden units. Orthogonal initialization is used for input-to-hidden weights, while hidden-to-hidden weights are initialized to identity matrices. The gradient norms are clipped to 1 during training. The log-magnitude of these sequences is fed into the models as the input data.
Appendix of Prophet Attention
CIDEr-c40, which is the default ranking score in the leaderboard, and rank the 1st. Compared with image captioning, the target of video captioning is the video clip, i.e., an ordered The dataset contain 10,000 video clips, and each video is paired with 20 annotated sentences. We use the official splits to report our results. CIDEr, which is built upon on n-gram matching, is used in our tests for performance evaluation. All re-implementations and our experiments were ran on V100 GPUs.
Appendix 1 0.1 Data augmentation
Figure 1 shows some examples of augmented MSCOCO images and captions.We perform image-3 Figure 1 illustrates the visual-textual alignment mechanisms of the three variants of our proposed SSRP. NL VR2 is a challenging visual reasoning task. During testing, we adopt beam search with a beam size of 5. We apply the same training and testing settings for Up-Down (Our Impl.) and SSRP Figure 3: Illustrations of the two image retrieval methods mentioned in our paper. Figure 4: Examples of generated relationships for different augmented images. Figure 5: Example of generated relationships for different augmented sentences.