Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation
Zhu, Wanrong, Wang, Xin, Fu, Tsu-Jui, Yan, An, Narayana, Pradyumna, Sone, Kazoo, Basu, Sugato, Wang, William Yang
–arXiv.org Artificial Intelligence
In the vision-and-language navigation (VLN) task, an agent follows natural language instructions and navigate in visual environments. Compared to the indoor navigation task that has been broadly studied, navigation in real-life outdoor environments remains a significant challenge with its complicated visual inputs and an insufficient amount of instructions that illustrate the intricate urban scenes. In this paper, we introduce a Multimodal Text Style Transfer (MTST) learning approach to mitigate the problem of data scarcity in outdoor navigation tasks by effectively leveraging external multimodal resources. We first enrich the navigation data by transferring the style of the instructions generated by Google Maps API, then pre-train the navigator with the augmented external outdoor navigation dataset. Experimental results show that our MTST learning approach is model-agnostic, and our MTST approach significantly outperforms the baseline models on the outdoor VLN task, improving task completion rate by 22\% relatively on the test set and achieving new state-of-the-art performance.
arXiv.org Artificial Intelligence
Jul-1-2020
- Country:
- North America > United States
- New York (0.04)
- California
- Santa Cruz County > Santa Cruz (0.04)
- Santa Barbara County > Santa Barbara (0.04)
- San Diego County > San Diego (0.04)
- Europe > Spain
- Catalonia > Barcelona Province > Barcelona (0.04)
- North America > United States
- Genre:
- Research Report > New Finding (0.34)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Natural Language (1.00)
- Machine Learning
- Neural Networks > Deep Learning (0.68)
- Statistical Learning (0.68)
- Information Technology > Artificial Intelligence