Visual Goal-Step Inference using wikiHow
Yang, Yue, Panagopoulou, Artemis, Lyu, Qing, Zhang, Li, Yatskar, Mark, Callison-Burch, Chris
–arXiv.org Artificial Intelligence
Procedural events can often be thought of as a high level goal composed of a sequence of steps. Inferring the sub-sequence of steps of a goal can help artificial intelligence systems reason about human activities. Past work in NLP has examined the task of goal-step inference for text. We introduce the visual analogue. We propose the Visual Goal-Step Inference (VGSI) task where a model is given a textual goal and must choose a plausible step towards that goal from among four candidate images. Our task is challenging for state-of-the-art muitimodal models. We introduce a novel dataset harvested from wikiHow that consists of 772,294 images representing human actions. We show that the knowledge Figure 1: An example Visual Goal-Step Inference learned from our data can effectively transfer Task: given a text goal (bake white fish), select the to other datasets like HowTo100M, increasing image (C) that represents a step towards that goal. the multiple-choice accuracy by 15% to 20%.
arXiv.org Artificial Intelligence
Apr-12-2021
- Country:
- Asia > China (0.04)
- North America > United States
- Pennsylvania (0.04)
- Genre:
- Workflow (0.54)
- Research Report (0.40)
- Industry:
- Education (0.67)
- Technology: