doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation
Roy, Parthib, Perisetla, Srinivasa, Shriram, Shashank, Krishnaswamy, Harsha, Keskar, Aryan, Greer, Ross
–arXiv.org Artificial Intelligence
Abstract--Human-interactive robotic systems, particularly autonomous vehicles (AVs), must effectively integrate human instructions into their motion planning. This paper introduces doScenes, a novel dataset designed to facilitate research on human-vehicle instruction interactions, focusing on short-term directives that directly influence vehicle motion. Unlike existing datasets that focus on ranking or scenelevel reasoning, doScenes emphasizes actionable directives tied to static and dynamic scene objects. This framework addresses limitations in prior research, such as reliance on simulated data or predefined action sets, by supporting nuanced and flexible responses in real-world scenarios. This work lays the foundation for developing learning strategies that seamlessly integrate human instructions into autonomous systems, advancing safe and effective human-vehicle collaboration. In the doScenes dataset, we augment each clip of temporal data with an instruction and a tag to indicate the instruction's referentiality.
arXiv.org Artificial Intelligence
Dec-8-2024
- Country:
- North America > United States
- California > Merced County > Merced (0.04)
- Asia
- North America > United States
- Genre:
- Research Report (0.40)
- Industry:
- Transportation > Ground > Road (1.00)
- Technology: