Technology
RadarOcc: Robust 3DOccupancy Prediction with 4DImaging Radar
Current methods predominantly rely on LiDAR or camera inputs for 3D occupancy prediction. These methods are susceptible to adverse weather conditions, limiting the all-weather deployment of self-driving cars. To improve perception robustness, we leverage the recent advances in automotive radars and introduce a novel approach that utilizes 4D imaging radar sensors for 3D occupancy prediction. Our method, RadarOcc, circumvents the limitations of sparse radar point clouds by directly processing the 4D radar tensor, thus preserving essential scene details. RadarOcc innovatively addresses the challenges associated with the voluminous and noisy 4D radar data by employing Doppler bins descriptors, sidelobe-aware spatial sparsification, and range-wise self-attention mechanisms. To minimize the interpolation errors associated with direct coordinate transformations, we also devise a spherical-based feature encoding followed by spherical-to-Cartesian feature aggregation. We benchmark various baseline methods based on distinct modalities on the public K-Radar dataset. The results demonstrate RadarOcc's state-of-the-art performance in radar-based 3D occupancy prediction and promising results even when compared with LiDARor camera-based methods. Additionally, we present qualitative evidence of the superior performance of 4D radar in adverse weather conditions and explore the impact of key pipeline components through ablation studies.
Operator Learning with Neural Fields: Tackling PDEs on General Geometries Supplemental Material Anonymous Author(s) Affiliation Address email
A.1 Initial Value Problem518 We use the datasets from Pfaff et al. (2021), and take the first and last frames of each trajectory as the519 input and output data for the initial value problem.520 Cylinder The dataset includes computational fluid dynamics (CFD) simulations of the flow around521 a cylinder, governed by the incompressible Navier-Stokes equation. These simulations were generated522 using COMSOL software, employing an irregular 2D-triangular mesh. The trajectory consists of 600523 timestamps, with a time interval of t =0 .01s between each timestamp.524 Airfoil The dataset contains CFD simulations of the flow around an airfoil, following the com-525 pressible Navier-Stokes equation. These simulations were conducted using SU2 software, using an526 irregular 2D-triangular mesh. The trajectory encompasses 600 timestamps, with a time interval of527 t =0 .008s between each timestamp.528 A.2 Dynamics Modeling529 2D-Navier-Stokes (Navier-Stokes) We consider the 2DNavier-Stokes equation as presented in Li530 et al. (2021); Yin et al. (2022).
Operator Learning with Neural Fields: Tackling PDEs on General Geometries
Machine learning approaches for solving partial differential equations require learning mappings between function spaces. While convolutional or graph neural networks are constrained to discretized functions, neural operators present a promising milestone toward mapping functions directly. Despite impressive results they still face challenges with respect to the domain geometry and typically rely on some form of discretization. In order to alleviate such limitations, we present CORAL, a new method that leverages coordinate-based networks for solving PDEs on general geometries. CORAL is designed to remove constraints on the input mesh, making it applicable to any spatial sampling and geometry. Its ability extends to diverse problem domains, including PDE solving, spatio-temporal forecasting, and geometry-aware inference. CORAL demonstrates robust performance across multiple resolutions and performs well in both convex and non-convex domains, surpassing or performing on par with state-of-the-art models.
Limitations
While our study identifies clear separations between model hypothesis classes, our best models still have not reached the consistency ceiling of the neural and behavioral benchmarks we have compared against. The latent future prediction dynamics modules of all the foundation models were pretrained on Physion just as the end-to-end models were, and those Physion trained dynamics modules were evaluated against neural and behavioral data, ultimately outperforming the end-to-end Physion dynamics. Despite our interest, pretraining the end-to-end models on datasets larger than Physion exceeds our current computational resources, as evidenced by models like FitVid requiring nearly a month of training on eight A100 GPUs with Physion alone. Therefore, the vision foundation models ultimately have to deal with the harder problem of generalizing to Physion compared to end-to-end models. While we believe our dynamically-equipped foundation model paradigm to be a generally promising way forward towards models with strong internal simulations, we identify in the Discussion ( 7), several ways that their encoder and dynamics modules can be improved, which we plan to explore in future work.
Masked Space-Time Hash Encoding for Efficient Dynamic Scene Reconstruction
In this paper, we propose the Masked Space-Time Hash encoding (MSTH), a novel method for efficiently reconstructing dynamic 3D scenes from multi-view or monocular videos. Based on the observation that dynamic scenes often contain substantial static areas that result in redundancy in storage and computations, MSTH represents a dynamic scene as a weighted combination of a 3D hash encodingOurs and a 4DPlenoptic Dataset 30 hash encoding.