DrivingScene: A Multi-Task Online Feed-Forward 3D Gaussian Splatting Method for Dynamic Driving Scenes
Hou, Qirui, Sun, Wenzhang, Zeng, Chang, Wang, Chunfeng, Li, Hao, Cui, Jianxun
–arXiv.org Artificial Intelligence
Modern autonomous vehicles are typically equipped with multiple cameras for 360-degree surround-view perception. Compared to fusion-based approaches that rely on multi-modal sensors like LiDAR or RaDAR[1, 2, 3], vision-only methods[4, 5] offer a more cost-effective and computationally efficient pathway for complex online perception tasks. However, reconstructing a large-scale, geometrically accurate, and photorealistic dynamic scene in real-time, solely from sparse and dynamic surround-view images, remains a significant and unresolved challenge. The pursuit of higher reconstruction fidelity has seen tremendous success with neural rendering techniques like NeRF [6] and 3DGS [7]. However, the majority of these methods, whether for static scenes like StreetGaussian [8], DrivingGaussian [9] or dynamic scenes like EmerNeRF [10], are bound by a per-scene optimization paradigm. This reliance on time-consuming offline training is incompatible with the real-time requirements of autonomous driving downstream tasks, necessitating a paradigm shift towards "feed-forward" reconstruction[11, 12, 13, 14]. This online approach has matured for static scenes, with methods like pixelSplat [15] and MVSplat [16] demonstrating its viability, and culminating in works like DrivingForward [17] which successfully handle sparse driving contexts. Y et, their foundational static world assumption inevitably leads to severe artifacts when confronted with moving vehicles.
arXiv.org Artificial Intelligence
Oct-30-2025
- Genre:
- Research Report (1.00)
- Industry:
- Information Technology (0.36)
- Technology:
- Information Technology > Artificial Intelligence
- Vision (1.00)
- Machine Learning (1.00)
- Robots > Autonomous Vehicles (0.56)
- Information Technology > Artificial Intelligence