Technology
A Neural Network Approach for Efficiently Answering Most Probable Explanation Queries in Probabilistic Models
We propose a novel neural networks based approach to efficiently answer arbitrary Most Probable Explanation (MPE) queries--a well-known NP-hard task--in large probabilistic models such as Bayesian and Markov networks, probabilistic circuits, and neural auto-regressive models. By arbitrary MPE queries, we mean that there is no predefined partition of variables into evidence and non-evidence variables. The key idea is to distill all MPE queries over a given probabilistic model into a neural network and then use the latter for answering queries, eliminating the need for time-consuming inference algorithms that operate directly on the probabilistic model. We improve upon this idea by incorporating inference-time optimization with self-supervised loss to iteratively improve the solutions and employ a teacher-student framework that provides a better initial network, which in turn, helps reduce the number of inference-time optimization steps. The teacher network utilizes a self-supervised loss function optimized for getting the exact MPE solution, while the student network learns from the teacher's near-optimal outputs through supervised loss. We demonstrate the efficacy and scalability of our approach on various datasets and a broad class of probabilistic models, showcasing its practical effectiveness.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
We explore how generating a chain of thought---a series of intermediate reasoning steps---significantly improves the ability of large language models to perform complex reasoning. In particular, we show how such reasoning abilities emerge naturally in sufficiently large language models via a simple method called chain of thought prompting, where a few chain of thought demonstrations are provided as exemplars in prompting. Experiments on three large language models show that chain of thought prompting improves performance on a range of arithmetic, commonsense, and symbolic reasoning tasks. The empirical gains can be striking. For instance, prompting a 540B-parameter language model with just eight chain of thought exemplars achieves state of the art accuracy on the GSM8K benchmark of math word problems, surpassing even finetuned GPT-3 with a verifier.
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
We study the sample complexity of learning an $\varepsilon$-optimal policy in an average-reward Markov decision process (MDP) under a generative model. For weakly communicating MDPs, we establish the complexity bound $\widetilde{O}\left(SA\frac{\mathsf{H}}{\varepsilon^2} \right)$, where $\mathsf{H}$ is the span of the bias function of the optimal policy and $SA$ is the cardinality of the state-action space. Our result is the first that is minimax optimal (up to log factors) in all parameters $S,A,\mathsf{H}$, and $\varepsilon$, improving on existing work that either assumes uniformly bounded mixing times for all policies or has suboptimal dependence on the parameters. We also initiate the study of sample complexity in general (multichain) average-reward MDPs.
Learning Spatially-Aware Language and Audio Embeddings
Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like the lion roar came from right behind me!. For a machine to have the same degree of comprehension, the machine must know what a lion is (semantic attribute), what the concept of behind is (spatial attribute) and how these pieces of linguistic information align with the semantic and spatial attributes of the sound (what a roar sounds like when its coming from behind). State-of-the-art audio foundation models, such as CLAP, which learn to map between audio scenes and natural textual descriptions, are trained on non-spatial audio and text pairs, and hence lack spatial awareness. In contrast, sound event localization and detection models are limited to recognizing sounds from a fixed number of classes, and they localize the source to absolute position (e.g., 0.2m) rather than a position described using natural language (e.g., next to me).
Game devs say Nvidia's DLSS 5 reveal blindsided them
PCWorld reports that Nvidia's DLSS 5 announcement caught major game developers from Ubisoft and Capcom off-guard, who were unaware their games would be featured in demonstrations. The generative AI technology faces significant backlash from gamers who criticize it as an "AI filter" that potentially devalues game aesthetics and may require two high-end GPUs. Despite being planned for fall 2026 release, DLSS 5 already raises concerns about artistic control and whether developers want this AI-enhanced visual processing in their games. Nvidia DLSS 5 is coming later this year, adding generative "AI" features to the performance-enhancing tech . Gamers are calling the tool an "Instagram yaas filter" and "AI slop," among other, less kind terms. The way that it adds detail to faces and seems to hijack -- or replace?
Rivian will provide 50,000 robotaxis to Uber in a deal worth 1.25 billion
Rivian will provide 50,000 robotaxis to Uber in a deal worth $1.25 billion Initial deployments will start in San Francisco and Miami. Rivian and Uber, with the former to provide the latter with 50,000 robotaxis in funding. This starts with Uber purchasing 10,000 Rivian R2 robotaxis, which will be deployed in San Francisco and Miami by 2028. If all goes well, Uber will scoop up 40,000 more robotaxis by 2030. The company plans to scale the initiative to 25 major cities by 2031.
Medieval chess was more inclusive than the world around it
Black, white, Muslim, or Christian: Players found common ground across the board. A black chess player about to win against a light-skinned cleric. Breakthroughs, discoveries, and DIY tips sent six days a week. Chess is widely seen as a great equalizer. Players from every social, racial, and economic class have squared off across the board for nearly 1,500 years, with victories determined solely by skill and strategy.
Essex police pause facial recognition camera use after study finds racial bias
Academics discover black people'significantly more likely' to be identified when compared with other ethnic groups Essex police have paused the use of live facial recognition (LFR) technology after a study found cameras were significantly more likely to target black people than people of other ethnicities. The move to suspend use of the AI-enabled systems was revealed by the Information Commissioner's Office (ICO), which regulates the use of the technology deployed so far by at least 13 police forces in London, south and north Wales, Leicestershire, Northamptonshire, Hampshire, Bedfordshire, Suffolk, Greater Manchester, West Yorkshire, Surrey and Sussex. The ICO said Essex police had paused LFR deployments "after identifying potential accuracy and bias risks" and warned other forces to have mitigations in place. LFR systems are either mounted to fixed locations or deployed in vans. In January, the home secretary, Shabana Mahmood, announced the number of LFR vans would increase five-fold, with 50 available to every police force in England and Wales. Essex commissioned University of Cambridge academics to conduct a study, which involved 188 actors walking past cameras being actively deployed from marked police vans in Chelmsford.
Windows 11's free video editor Clipchamp now requires OneDrive
PCWorld reports that Microsoft's Clipchamp video editor in Windows 11 now mandates OneDrive for saving and editing video projects. This change significantly impacts users who prefer local storage, as locally saved projects become uneditable archives that cannot be modified. New Clipchamp projects automatically sync to OneDrive accounts, though media files within projects may not always require cloud synchronization. Microsoft is changing how Clipchamp--the built-in free video editor for Windows 11--works. The program now requires video projects to be saved to Microsoft's OneDrive cloud storage service in order to continue editing them, reports Windows Latest .
Crimson Desert: The all-you-can-eat video game divides critics
Video game fans and big, blockbuster releases have had an uneasy relationship in recent years. As so-called triple-A games get more expensive to make, the publishers behind them are accused of taking fewer risks and failing to try new things. But highly anticipated new release Crimson Desert asks a different question - what if a big-budget, graphically advanced game tried to do absolutely everything? The ambitious action-adventure's been compared to a buffet, presenting players with a smorgasbord of ideas, gameplay styles and quests to gorge on. While some have praised it as a feast, others have found it overstuffed, with some undercooked morsels behind the impressive presentation.