S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
Xie, Yichen, Xu, Runsheng, He, Tong, Hwang, Jyh-Jing, Luo, Katie, Ji, Jingwei, Lin, Hubert, Chen, Letian, Lu, Yiren, Leng, Zhaoqi, Anguelov, Dragomir, Tan, Mingxing
–arXiv.org Artificial Intelligence
The latest advancements in multi-modal large language models (MLLMs) have spurred a strong renewed interest in end-to-end motion planning approaches for autonomous driving. Many end-to-end approaches rely on human annotations to learn intermediate perception and prediction tasks, while purely self-supervised approaches--which directly learn from sensor inputs to generate planning trajectories without human annotations--often underperform the state of the art. W e observe a key gap in the input representation space: end-to-end approaches built on MLLMs are often pretrained with reasoning tasks in 2D image space rather than the native 3D space in which autonomous vehicles plan. T o this end, we propose S4-Driver, a s calable s elf-s upervised motion planning algorithm with s patio-temporal visual representation, based on the popular PaLI [9] multimodal large language model. S4-Driver uses a novel sparse volume strategy to seamlessly transform the strong visual representation of MLLMs from perspective view to 3D space without the need to finetune the vision encoder . This representation aggregates multi-view and multi-frame visual inputs and enables better prediction of planning trajectories in 3D space. T o validate our method, we run experiments on both nuScenes and W aymo Open Motion Dataset (with in-house camera data). Results show that S4-Driver performs favorably against existing supervised multi-task approaches while requiring no human annotations. It also demonstrates great scalability when pre-trained on large volumes of unannotated driving logs.
arXiv.org Artificial Intelligence
Jun-4-2025
- Genre:
- Research Report (0.84)
- Industry:
- Information Technology (0.50)
- Transportation > Ground
- Road (0.68)
- Technology: