multimodal coreference resolution
UNITER-Based Situated Coreference Resolution with Rich Multimodal Input
Huang, Yichen, Wang, Yuchen, Tam, Yik-Cheung
We propose a UNITER(Chen et al. 2020)-based model for The goal of Situated and Interactive Multimodal Conversation SIMMC 2.0. UNITER is proposed in computer vision (CV) (SIMMC) 2.0 (Kottur et al. 2021) is to aid the conversational for universal embeddings for image and text. To achieve this AI community in developing successful multimodal goal, UNITER is pre-trained with masked language modelling, assistant agents capable of handling real-world multimodal masked region modelling and word-region alignment dialog inputs.