Hint-AD: Holistically Aligned Interpretability in End-to-End Autonomous Driving
Ding, Kairui, Chen, Boyuan, Su, Yuchen, Gao, Huan-ang, Jin, Bu, Sima, Chonghao, Zhang, Wuqiang, Li, Xiaohui, Barsch, Paul, Li, Hongyang, Zhao, Hao
–arXiv.org Artificial Intelligence
Human-friendly natural language has been explored for tasks such as driving explanation and 3D captioning. However, previous works primarily focused on the paradigm of declarative interpretability, where the natural language interpretations are not grounded in the intermediate outputs of AD systems, making the interpretations only declarative. In contrast, aligned interpretability establishes a connection between language and the intermediate outputs of AD systems. Here we introduce Hint-AD, an integrated AD-language system that generates language aligned with the holistic perceptionprediction-planning outputs of the AD model. By incorporating the intermediate outputs and a holistic token mixer sub-network for effective feature adaptation, Hint-AD achieves desirable accuracy, achieving state-of-the-art results in driving language tasks including driving explanation, 3D dense captioning, and command prediction. To facilitate further study on driving explanation task on nuScenes, we also introduce a human-labeled dataset, Nu-X.
arXiv.org Artificial Intelligence
Sep-10-2024
- Country:
- Asia
- Middle East > Republic of Türkiye
- Karaman Province > Karaman (0.04)
- China > Shanghai
- Shanghai (0.04)
- Middle East > Republic of Türkiye
- Asia
- Genre:
- Research Report (0.82)
- Industry:
- Automobiles & Trucks (1.00)
- Information Technology (0.84)
- Transportation > Ground
- Road (1.00)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Robots > Autonomous Vehicles (1.00)
- Representation & Reasoning (1.00)
- Natural Language > Large Language Model (1.00)
- Machine Learning > Neural Networks
- Deep Learning (0.95)
- Information Technology > Artificial Intelligence