Generalized Hadamard-Product Fusion Operators for Visual Question Answering
Duke, Brendan, Taylor, Graham W.
Multimodal applications offer a challenge to model selection in machine learning, as the interactions between different data modalities (e.g., between video and audio, or between images and questions) may require a complicated prior in order to accurately capture regularities necessary for downstream tasks. The particular multimodal application that we consider in this work is that of visual question-answering (VQA), i.e., of producing a natural language response to the combined input consisting of an image and a natural language question pertaining specifically to that image. In the case of VQA, the complexity of the task exists in both extracting useful feature representations from the question and the image, as well as the "fusion" of these feature representations by combining them in order to predict the answer to the question. In this paper, we focus on the problem of combining feature representations in the VQA task through models based on a generic "fusion operator" definition. We illustrate the complexity of model selection for the VQA task by designing a class of multimodal fusion operators, each of which combines raw question and visual data streams to predict answers to the given questions based on an image. We evaluate specific instances of high performing fusion operators belonging to the same design class. We evaluate and discuss three multimodal fusion architectural components that emerge as improving performance as part of the investigation into the general class of fusion operators: 1) The use of a gating mechanism, wherein individual features extracted by the fusion operator are turned on or off by a multiplicative interaction.
Mar-25-2018