IWISDM: Assessing instruction following in multimodal models at scale

Lei, Xiaoxuan, Gomez, Lucas, Bai, Hao Yuan, Bashivan, Pouya

arXiv.org Artificial Intelligence 

The ability to perform complex tasks from detailed instructions is a key to many remarkable achievements of our species. As humans, we are not only capable of performing a wide variety of tasks but also very complex ones that may entail hundreds or thousands of steps to complete. Large language models and their more recent multimodal counterparts that integrate textual and visual inputs have achieved unprecedented success in performing complex tasks. Yet, most existing benchmarks are largely confined to single-modality inputs (either text or vision), narrowing the scope of multimodal assessments, particularly for instruction-following in multimodal contexts. To bridge this gap, we introduce the instructed-Virtual VISual Decision Making (iWISDM) environment engineered to generate a limitless array of vision-language tasks of varying complexity. Using iWISDM, we compiled three distinct benchmarks of instruction following visual tasks across varying complexity levels and evaluated several newly developed multimodal models on these benchmarks. Our findings establish iWISDM as a robust benchmark for assessing the instructional adherence of both existing and emergent multimodal models and highlight a large gap between these models' ability to precisely follow instructions with that of humans. A typical day in most people's lives involves hundreds or thousands of tasks. Most of which are performed without explicit attention. Just in between getting up and getting to work, one may have already performed 5-15 tasks (taking a shower, shaving, making coffee, getting dressed, etc.). Teaching artificial agents to perform similar seemingly mundane tasks has proven to be an extremely difficult computational problem (Konar, 2018). The challenge becomes more apparent when one realizes that each of these seemingly mundane tasks such as making coffee involves tens of steps/actions (Figure 1). The challenge becomes increasingly more significant once we consider more complex tasks such as operating a device or assembling a piece of furniture from its instruction manual. And yet, these tasks are performed proficiently by most individuals in most situations. Large Language Models (LLMs) have become increasingly capable of comprehending natural language across wide topics and contexts, allowing them to hold meaningful conversations, give expert advice, and analyze data among other features (Brown et al., 2020; Ouyang et al., 2022; Radford et al., 2019). In the meantime, their multimodal counterparts are starting to emerge, signalling broader application of such models across industries. Large Multimodal Models (LMMs) are generally capable of receiving and responding in a range of possible modalities including visual, text, and audio (Alayrac et al., 2022; Liu et al., 2023b; Achiam et al., 2023). These authors contributed equally to this work.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found