PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs

Zhang, Zixin, Chen, Kanghao, Lin, Xingwang, Jiang, Lutao, Zheng, Xu, Lyu, Yuanhuiyi, Guo, Litao, Li, Yinchuan, Chen, Ying-Cong

arXiv.org Artificial Intelligence 

Knowin Figure 1: For an Embodied Agent, using physical tools is crucial in many tasks. Phys-ToolBench (Bottom) systematically evaluates the understanding of physical tools of multimodal LLMs. The benchmark is designed with three progressive levels of difficulty and employs a Visual Question Answering (VQA) format. Notice that in the actual benchmark, tools in the images are numerically labeled, and images here are for illustrative purposes only. The ability to use, understand, and create tools is a hallmark of human intelligence, enabling sophisticated interaction with the physical world. For any general-purpose intelligent agent to achieve true versatility, it must also master these fundamental skills. While modern Multimodal Large Language Models (MLLMs) leverage their extensive common knowledge for high-level planning in embodied AI and in downstream Vision-Language-Action (VLA) models, the extent of their true understanding of physical tools remains unquantified. To bridge this gap, we present PhysT oolBench, the first benchmark dedicated to evaluating the comprehension of physical tools by MLLMs. Our benchmark is structured as a Visual Question Answering (VQA) dataset comprising over 1,000 image-text pairs. Our comprehensive evaluation of 32 MLLMs--spanning proprietary, open-source, specialized embodied, and backbones in VLAs--reveals a significant deficiency in the tool understanding.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found