UNITER-Based Situated Coreference Resolution with Rich Multimodal Input

Huang, Yichen, Wang, Yuchen, Tam, Yik-Cheung

Dec-7-2021–arXiv.org Artificial Intelligence

We propose a UNITER(Chen et al. 2020)-based model for The goal of Situated and Interactive Multimodal Conversation SIMMC 2.0. UNITER is proposed in computer vision (CV) (SIMMC) 2.0 (Kottur et al. 2021) is to aid the conversational for universal embeddings for image and text. To achieve this AI community in developing successful multimodal goal, UNITER is pre-trained with masked language modelling, assistant agents capable of handling real-world multimodal masked region modelling and word-region alignment dialog inputs.

multimodal coreference resolution, scene graph, uniter, (8 more...)

arXiv.org Artificial Intelligence

Dec-7-2021

arXiv.org PDF

Add feedback

Country:
- North America > United States
  - New York (0.04)
  - Minnesota > Hennepin County
    - Minneapolis (0.04)
  - California > San Diego County
    - San Diego (0.04)
- Asia > China
  - Shanghai > Shanghai (0.04)

Genre:
- Research Report (0.50)

Technology:
- Information Technology > Artificial Intelligence
  - Vision (1.00)
  - Machine Learning (1.00)
  - Natural Language > Discourse & Dialogue (0.71)