doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation

Roy, Parthib, Perisetla, Srinivasa, Shriram, Shashank, Krishnaswamy, Harsha, Keskar, Aryan, Greer, Ross

arXiv.org Artificial Intelligence 

Abstract--Human-interactive robotic systems, particularly autonomous vehicles (AVs), must effectively integrate human instructions into their motion planning. This paper introduces doScenes, a novel dataset designed to facilitate research on human-vehicle instruction interactions, focusing on short-term directives that directly influence vehicle motion. Unlike existing datasets that focus on ranking or scenelevel reasoning, doScenes emphasizes actionable directives tied to static and dynamic scene objects. This framework addresses limitations in prior research, such as reliance on simulated data or predefined action sets, by supporting nuanced and flexible responses in real-world scenarios. This work lays the foundation for developing learning strategies that seamlessly integrate human instructions into autonomous systems, advancing safe and effective human-vehicle collaboration. In the doScenes dataset, we augment each clip of temporal data with an instruction and a tag to indicate the instruction's referentiality.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found