BEV-Locator: An End-to-end Visual Semantic Localization Network Using Multi-View Images

Zhang, Zhihuang, Xu, Meng, Zhou, Wenqiang, Peng, Tao, Li, Liang, Poslad, Stefan

Nov-27-2022–arXiv.org Artificial Intelligence

Accurate localization ability is fundamental in autonomous driving. Traditional visual localization frameworks approach the semantic map-matching problem with geometric models, which rely on complex parameter tuning and thus hinder large-scale deployment. In this paper, we propose BEV-Locator: an end-to-end visual semantic localization neural network using multi-view camera images. Specifically, a visual BEV (Birds-Eye-View) encoder extracts and flattens the multi-view images into BEV space. While the semantic map features are structurally embedded as map queries sequence. Then a cross-model transformer associates the BEV features and semantic map queries. The localization information of ego-car is recursively queried out by cross-attention modules. Finally, the ego pose can be inferred by decoding the transformer outputs. We evaluate the proposed method in large-scale nuScenes and Qcraft datasets. The experimental results show that the BEV-locator is capable to estimate the vehicle poses under versatile scenarios, which effectively associates the cross-model information from multi-view images and global semantic maps. The experiments report satisfactory accuracy with mean absolute errors of 0.052m, 0.135m and 0.251$^\circ$ in lateral, longitudinal translation and heading angle degree.

localization, machine learning, natural language, (18 more...)

arXiv.org Artificial Intelligence

Nov-27-2022

arXiv.org PDF

Add feedback

Country:
- Asia > China (0.29)
- Europe > Germany (0.28)

Genre:
- Research Report > New Finding (0.34)

Industry:
- Transportation
  - Ground > Road (0.89)
  - Infrastructure & Services (0.68)

Technology:
- Information Technology
  - Artificial Intelligence
    - Machine Learning > Neural Networks
      - Deep Learning (0.46)
      - Perceptrons (0.46)
    - Natural Language (0.93)
    - Representation & Reasoning (1.00)
    - Robots > Autonomous Vehicles (0.89)
    - Vision (1.00)
  - Sensing and Signal Processing (1.00)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found