CMU, Oxford & Facebook Cross-Lingual Vision-Language Model Achieves New SOTA in Zero-Shot Setting
Building versatile vision-language models that not only work on a single language but can generalize across all the world's approximately 7,000 languages is difficult -- and the task becomes even more challenging if the model is transferred without any additional annotated training data. To tackle this issue, a research team from Carnegie Mellon, Oxford and Facebook AI has proposed a transformer-based model, Multilingual Multimodal Pretraining (MMP), that can learn contextualized multilingual multimodal embeddings under a zero-shot setting. Recent research in cross-lingual transfer learning has demonstrated that models using only English annotation can nonetheless generalize to a non-English language. This success is attributed to the shared underlying vocabulary or structure amongst many languages. For example, many English and German words stem from the same origin, and many languages have the same recursive structures.
Dec-13-2021, 00:09:42 GMT
- Technology: