geiger
Scalable Offline Metrics for Autonomous Driving
Aich, Animikh, Kulkarni, Adwait, Ohn-Bar, Eshed
Real-world evaluation of perception-based planning models for robotic systems, such as autonomous vehicles, can be safely and inexpensively conducted offline, i.e. by computing model prediction error over a pre-collected validation dataset with ground-truth annotations. However, extrapolating from offline model performance to online settings remains a challenge. In these settings, seemingly minor errors can compound and result in test-time infractions or collisions. This relationship is understudied, particularly across diverse closed-loop metrics and complex urban maneuvers. In this work, we revisit this undervalued question in policy evaluation through an extensive set of experiments across diverse conditions and metrics. Based on analysis in simulation, we find an even worse correlation between offline and online settings than reported by prior studies, casting doubts on the validity of current evaluation practices and metrics for driving policies. Next, we bridge the gap between offline and online evaluation. We investigate an offline metric based on epistemic uncertainty, which aims to capture events that are likely to cause errors in closed-loop settings. The resulting metric achieves over 13% improvement in correlation compared to previous offline metrics. We further validate the generalization of our findings beyond the simulation environment in real-world settings, where even greater gains are observed.
Pseudo-Simulation for Autonomous Driving
Cao, Wei, Hallgarten, Marcel, Li, Tianyu, Dauner, Daniel, Gu, Xunjiang, Wang, Caojun, Miron, Yakov, Aiello, Marco, Li, Hongyang, Gilitschenski, Igor, Ivanovic, Boris, Pavone, Marco, Geiger, Andreas, Chitta, Kashyap
Existing evaluation paradigms for Autonomous Vehicles (AVs) face critical limitations. Real-world evaluation is often challenging due to safety concerns and a lack of reproducibility, whereas closed-loop simulation can face insufficient realism or high computational costs. Open-loop evaluation, while being efficient and data-driven, relies on metrics that generally overlook compounding errors. In this paper, we propose pseudo-simulation, a novel paradigm that addresses these limitations. Pseudo-simulation operates on real datasets, similar to open-loop evaluation, but augments them with synthetic observations generated prior to evaluation using 3D Gaussian Splatting. Our key idea is to approximate potential future states the AV might encounter by generating a diverse set of observations that vary in position, heading, and speed. Our method then assigns a higher importance to synthetic observations that best match the AV's likely behavior using a novel proximity-based weighting scheme. This enables evaluating error recovery and the mitigation of causal confusion, as in closed-loop benchmarks, without requiring sequential interactive simulation. We show that pseudo-simulation is better correlated with closed-loop simulations ($R^2=0.8$) than the best existing open-loop approach ($R^2=0.7$). We also establish a public leaderboard for the community to benchmark new methodologies with pseudo-simulation. Our code is available at https://github.com/autonomousvision/navsim.
Why is deep sleep so important to memory? It's about time.
It's no hidden health secret that sleep is really good for us. It helps our immune systems and supports almost every organ system in the body. We've also known for almost two decades that the slow, synchronous electrical waves in the brain during deep sleep supports memory formation. However, we did not know exactly how the brain does this until now. These slow waves make the neocortexโwhere long-term memory is stored in the brainโparticularly receptive to new information.
LaRa: Efficient Large-Baseline Radiance Fields
Chen, Anpei, Xu, Haofei, Esposito, Stefano, Tang, Siyu, Geiger, Andreas
Radiance field methods have achieved photorealistic novel view synthesis and geometry reconstruction. But they are mostly applied in per-scene optimization or small-baseline settings. While several recent works investigate feed-forward reconstruction with large baselines by utilizing transformers, they all operate with a standard global attention mechanism and hence ignore the local nature of 3D reconstruction. We propose a method that unifies local and global reasoning in transformer layers, resulting in improved quality and faster convergence. Our model represents scenes as Gaussian Volumes and combines this with an image encoder and Group Attention Layers for efficient feed-forward reconstruction. Experimental results demonstrate that our model, trained for two days on four GPUs, demonstrates high fidelity in reconstructing 360 deg radiance fields, and robustness to zero-shot and out-of-domain testing. Our project Page: https://apchenstu.github.io/LaRa/.
StyleNeRF: A Style-based 3D-Aware Generator for High-resolution Image Synthesis
Gu, Jiatao, Liu, Lingjie, Wang, Peng, Theobalt, Christian
We propose StyleNeRF, a 3D-aware generative model for photo-realistic high-resolution image synthesis with high multi-view consistency, which can be trained on unstructured 2D images. Existing approaches either cannot synthesize high-resolution images with fine details or yield noticeable 3D-inconsistent artifacts. In addition, many of them lack control over style attributes and explicit 3D camera poses. StyleNeRF integrates the neural radiance field (NeRF) into a style-based generator to tackle the aforementioned challenges, i.e., improving rendering efficiency and 3D consistency for high-resolution image generation. We perform volume rendering only to produce a low-resolution feature map and progressively apply upsampling in 2D to address the first issue. To mitigate the inconsistencies caused by 2D upsampling, we propose multiple designs, including a better upsampler and a new regularization loss. With these designs, StyleNeRF can synthesize high-resolution images at interactive rates while preserving 3D consistency at high quality. StyleNeRF also enables control of camera poses and different levels of styles, which can generalize to unseen views. It also supports challenging tasks, including zoom-in and-out, style mixing, inversion, and semantic editing.
The Success of Conversational AI and the AI Evaluation Challenge it Reveals
Research interest in Conversational AI has experienced a massive growth over the last few years and several recent advancements have enabled systems to produce rich and varied turns in conversations similar to humans. However, this apparent creativity is also creating a real challenge in the objective evaluation of such systems as authors are becoming reliant on crowd worker opinions as the primary measurement of success and, so far, few papers are reporting all that is necessary for others to compare against in their own crowd experiments. This challenge is not unique to ConvAI, but demonstrates as AI systems mature in more "human" tasks that involve creativity and variation, evaluation strategies need to mature with them. Conversational AI, or ConvAI as it has been abbreviated, is a sub-field of artificial intelligence (AI) where the goal is to build an autonomous agent that is capable of maintaining natural discourse with a human over some interface such as text or speech. The purpose may be to help humans perform tasks as a virtual/digital assistant, provide a natural language interface to another system as in information retrieval or navigation systems, or simply to converse like one would with an open domain chatbot.
Solving the Torpedo Scheduling Problem
Geiger, Martin Josef, Kletzander, Lucas, Musliu, Nysret
The article presents a solution approach for the Torpedo Scheduling Problem, an operational planning problem found in steel production. The problem consists of the integrated scheduling and routing of torpedo cars, i. e. steel transporting vehicles, from a blast furnace to steel converters. In the continuous metallurgic transformation of iron into steel, the discrete transportation step of molten iron must be planned with considerable care in order to ensure a continuous material flow. The problem is solved by a Simulated Annealing algorithm, coupled with an approach of reducing the set of feasible material assignments. The latter is based on logical reductions and lower bound calculations on the number of torpedo cars. Experimental investigations are performed on a larger number of problem instances, which stem from the 2016 implementation challenge of the Association of Constraint Programming (ACP). Our approach was ranked first (joint first place) in the 2016 ACP challenge and found optimal solutions for all used instances in this challenge.
Owning Guns Is Sort of Like Owning Rattlesnakes
In his short story "Rattlesnakes and Men," science fiction author Michael Bishop describes a town where everyone is required by law to own a dangerous rattlesnake. It's a scenario that he says is no more absurd than how America treats access to guns. "We lost our son at Virginia Tech in 2007, in the shootings there," Bishop says in Episode 322 of the Geek's Guide to the Galaxy podcast. "I had been opposed to the laxity of our gun laws for a long, long time, and that just hardened both my wife and me on that particular point." The story features an organization called the Nokuse Rattlesnake Alliance, which forces schools to adopt living pit vipers, spends large sums to corrupt local politicians, and hides the truth about the number of snakebite victims.
UnFlow: Unsupervised Learning of Optical Flow With a Bidirectional Census Loss
Meister, Simon (TU Darmstadt) | Hur, Junhwa (TU Darmstadt) | Roth, Stefan (TU Darmstadt)
In the era of end-to-end deep learning, many advances in computer vision are driven by large amounts of labeled data. In the optical flow setting, however, obtaining dense per-pixel ground truth for real scenes is difficult and thus such data is rare. Therefore, recent end-to-end convolutional networks for optical flow rely on synthetic datasets for supervision, but the domain mismatch between training and test scenarios continues to be a challenge. Inspired by classical energy-based optical flow methods, we design an unsupervised loss based on occlusion-aware bidirectional flow estimation and the robust census transform to circumvent the need for ground truth flow. On the KITTI benchmarks, our unsupervised approach outperforms previous unsupervised deep networks by a large margin, and is even more accurate than similar supervised methods trained on synthetic datasets alone. By optionally fine-tuning on the KITTI training data, our method achieves competitive optical flow accuracy on the KITTI 2012 and 2015 benchmarks, thus in addition enabling generic pre-training of supervised networks for datasets with limited amounts of ground truth.