GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
Deshpande, Abhay, Deng, Yuquan, Ray, Arijit, Salvador, Jordi, Han, Winson, Duan, Jiafei, Zeng, Kuo-Hao, Zhu, Yuke, Krishna, Ranjay, Hendrix, Rose
–arXiv.org Artificial Intelligence
GraspMolmo predicts semantically appropriate, stable grasps conditioned on a natural language instruction and a single RGB-D frame. For instance, given "pour me some tea," GraspMolmo selects a grasp on a teapot handle rather than its body. Unlike prior TOG methods, which are limited by small datasets, simplistic language, and unrealistically simple scenes, GraspMolmo learns from PRISM, a novel large-scale synthetic dataset of 379k samples featuring complex environments and diverse, realistic task descriptions. We fine-tune the Molmo vision-language model on this data, enabling GraspMolmo to generalize to novel open-vocabulary instructions and objects. In challenging real-world evaluations, GraspMolmo achieves state-of-the-art results, with a 70% prediction success on complex tasks, compared to the 35% achieved by the next best alternative. GraspMolmo also successfully demonstrates the ability to predict semantically correct bimanual grasps zero-shot. We release our synthetic dataset, code, model, and benchmarks to accelerate research in task-semantic robotic manipulation, which, along with videos, are available at this URL.
arXiv.org Artificial Intelligence
Sep-16-2025
- Country:
- North America > United States (0.67)
- Genre:
- Research Report (0.82)
- Industry:
- Leisure & Entertainment > Games > Computer Games (0.46)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Robots (1.00)
- Representation & Reasoning (1.00)
- Natural Language > Large Language Model (1.00)
- Machine Learning > Neural Networks
- Deep Learning (0.68)
- Information Technology > Artificial Intelligence