Goto

Collaborating Authors

 Energy


Model-Based Control with Sparse Neural Dynamics

Neural Information Processing Systems

Learning predictive models from observations using deep neural networks (DNNs) is a promising new approach to many real-world planning and control problems. However, common DNNs are too unstructured for effective planning, and current control methods typically rely on extensive sampling or local gradient descent. In this paper, we propose a new framework for integrated model learning and predictive control that is amenable to efficient optimization algorithms. Specifically, we start with a ReLU neural model of the system dynamics and, with minimal losses in prediction accuracy, we gradually sparsify it by removing redundant neurons. This discrete sparsification process is approximated as a continuous problem, enabling an end-to-end optimization of both the model architecture and the weight parameters. The sparsified model is subsequently used by a mixed-integer predictive controller, which represents the neuron activations as binary variables and employs efficient branch-and-bound algorithms. Our framework is applicable to a wide variety of DNNs, from simple multilayer perceptrons to complex graph neural dynamics. It can efficiently handle tasks involving complicated contact dynamics, such as object pushing, compositional object sorting, and manipulation of deformable objects. Numerical and hardware experiments show that, despite the aggressive sparsification, our framework can deliver better closed-loop performance than existing state-of-the-art methods.


CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders

Neural Information Processing Systems

A vital and rapidly growing application, remote sensing offers vast yet sparsely labeled, spatially aligned multimodal data; this makes self-supervised learning algorithms invaluable. We present CROMA: a framework that combines contrastive and reconstruction self-supervised objectives to learn rich unimodal and multimodal representations. Our method separately encodes masked-out multispectral optical and synthetic aperture radar samples--aligned in space and time--and performs cross-modal contrastive learning. Another encoder fuses these sensors, producing joint multimodal encodings that are used to predict the masked patches via a lightweight decoder. We show that these objectives are complementary when leveraged on spatially aligned multimodal data. We also introduce X-and 2D-ALiBi, which spatially biases our cross-and self-attention matrices. These strategies improve representations and allow our models to effectively extrapolate to images up to $17.6\times$ larger at test-time.


Improving *day-ahead* Solar Irradiance Time Series Forecasting by Leveraging Spatio-Temporal Context

Neural Information Processing Systems

Nonetheless, the inherent variability of solar irradiance poses a significant challenge for seamlessly integrating solar power into the electrical grid. While the majority of prior research has centered on employing purely time series-based methodologies for solar forecasting, only a limited number of studies have taken into account factors such as cloud cover or the surrounding physical context.In this paper, we put forth a deep learning architecture designed to harness spatio-temporal context using satellite data, to attain highly accurate day-ahead time-series forecasting for any given station, with a particular emphasis on forecasting Global Horizontal Irradiance (GHI). We also suggest a methodology to extract a distribution for each time step prediction, which can serve as a very valuable measure of uncertainty attached to the forecast. When evaluating models, we propose a testing scheme in which we separate particularly difficult examples from easy ones, in order to capture the model performances in crucial situations, which in the case of this study are the days suffering from varying cloudy conditions. Furthermore, we present a new multi-modal dataset gathering satellite imagery over a large zone and time series for solar irradiance and other related physical variables from multiple geographically diverse solar stations. Our approach exhibits robust performance in solar irradiance forecasting, including zero-shot generalization tests at unobserved solar stations, and holds great promise in promoting the effective integration of solar power into the grid.


5 incredible aerospace breakthroughs in 2025

Popular Science

The Vera C. Rubin Observatory won our Innovation of the Year honors. Breakthroughs, discoveries, and DIY tips sent every weekday. From the most detailed movie of the night sky ever made to the first commercial soft landing on the moon, this year has been an inflection point for exploring and understanding the vast expanse above our heads. We also saw breakthroughs in small changes to commercial airliners that improve efficiency, as well as a new type of rocket engine that might be the future of extremely high speed air travel, plus the closest view of Mercury we've ever seen! Vera C. Rubin Observatory by U.S. National Science Foundation & Department of Energy: World's largest digital camera to conduct 10-year survey of the night sky Prepare to see space like never before.


The 5 coolest gadget innovations of 2025

Popular Science

We may earn revenue from the products available on this page and participate in affiliate programs. Deep down, we want to be cyborgs. We spend huge chunks of time interacting with technology every day, but the friction created by devices and interfaces persists. This year, we got closer than we have been to tech that truly augments reality. Meta took its smart glasses beyond its beginning as a simple content creation tool.


Provable Subspace Identification Under Post-Nonlinear Mixtures

Neural Information Processing Systems

Unsupervised mixture learning (UML) aims at identifying linearly or nonlinearly mixed latent components in a blind manner. UML is known to be challenging: Even learning linear mixtures requires highly nontrivial analytical tools, e.g., independent component analysis or nonnegative matrix factorization. In this work, the post-nonlinear (PNL) mixture model---where {\it unknown} element-wise nonlinear functions are imposed onto a linear mixture---is revisited. The PNL model is widely employed in different fields ranging from brain signal classification, speech separation, remote sensing, to causal discovery. To identify and remove the unknown nonlinear functions, existing works often assume different properties on the latent components (e.g., statistical independence or probability-simplex structures).


SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery

Neural Information Processing Systems

Unsupervised pre-training methods for large vision models have shown to enhance performance on downstream supervised tasks. Developing similar techniques for satellite imagery presents significant opportunities as unlabelled data is plentiful and the inherent temporal and multi-spectral structure provides avenues to further improve existing pre-training strategies. In this paper, we present SatMAE, a pre-training framework for temporal or multi-spectral satellite imagery based on Masked Autoencoder (MAE). To leverage temporal information, we include a temporal embedding along with independently masking image patches across time. In addition, we demonstrate that encoding multi-spectral data as groups of bands with distinct spectral positional encodings is beneficial. Our approach yields strong improvements over previous state-of-the-art techniques, both in terms of supervised learning performance on benchmark datasets (up to $\uparrow$ 7%), and transfer learning performance on downstream remote sensing tasks, including land cover classification (up to $\uparrow$ 14%) and semantic segmentation.


Cozy up (safely) to an e-scooter's lithium battery yule log

Popular Science

Breakthroughs, discoveries, and DIY tips sent every weekday. The United States Consumer Product Safety Commission (CPSC) is well known for getting their point across on social media. A seven-minute montage of mannequins succumbing to 4th of July firework injuries may be an unconventional way to warn about the dangers of recreational explosives--but try forgetting those images when lighting your next bottle rocket. In similar pyrotechnic fashion, the CPSC is warning everyone to take extra care during the holidays when it comes to all kinds of combustible, seasonally appropriate objects. On December 22, the commission illustrated how some gifts are far more flammable than others with its 30-minute Escooter Lithium-Ion Battery Yule Log video.


The showers and baths keeping data centre tech cool

BBC News

They work 24/7 at high speeds and get searingly hot - but data centre computer chips get plenty of pampering. Some of them basically live at the spa. We'll have fluid that comes up and [then] shower down, or trickle down, onto a component, says Jonathan Ballon, chief executive at liquid cooling firm Iceotope. Some things will get sprayed. In other cases, the industrious gizmos recline in circulating baths of fluid, which ferries away the heat they generate, enabling them to function at very high speeds, known as overclocking.


Deep Learning for Primordial $B$-mode Extraction

arXiv.org Machine Learning

The search for primordial gravitational waves is a central goal of cosmic microwave background (CMB) surveys. Isolating the characteristic $B$-mode polarization signal sourced by primordial gravitational waves is challenging for several reasons: the amplitude of the signal is inherently small; astrophysical foregrounds produce $B$-mode polarization contaminating the signal; and secondary $B$-mode polarization fluctuations are produced via the conversion of $E$ modes. Current and future low-noise, multi-frequency observations enable sufficient precision to address the first two of these challenges such that secondary $B$ modes will become the bottleneck for improved constraints on the amplitude of primordial gravitational waves. The dominant source of secondary $B$-mode polarization is gravitational lensing by large scale structure. Various strategies have been developed to estimate the lensing deflection and to reverse its effects the CMB, thus reducing confusion from lensing $B$ modes in the search for primordial gravitational waves. However, a few complications remain. First, there may be additional sources of secondary $B$-mode polarization, for example from patchy reionization or from cosmic polarization rotation. Second, the statistics of delensed CMB maps can become complicated and non-Gaussian, especially when advanced lensing reconstruction techniques are applied. We previously demonstrated how a deep learning network, ResUNet-CMB, can provide nearly optimal simultaneous estimates of multiple sources of secondary $B$-mode polarization. In this paper, we show how deep learning can be applied to estimate and remove multiple sources of secondary $B$-mode polarization, and we further show how this technique can be used in a likelihood analysis to produce nearly optimal, unbiased estimates of the amplitude of primordial gravitational waves.