Inourexperiments, we visualize images for the mini-ImageNet dataset, and thus we utilize the image generator from the work [10] asG, which is pre-trained on ImageNet, and we pre-trainf on the mini-ImageNet dataset.
Joint visual and language modeling on large-scale datasets has recently shown good progress in multi-modal tasks when compared to single modal learning.