Instructional Material
Fire360: A Benchmark for Robust Perception and Episodic Memory in Degraded 360 Firefighting Video
Modern AI systems struggle most in environments where reliability is critical - scenes with smoke, poor visibility, and structural deformation. Each year, tens of thousands of firefighters are injured on duty, often due to breakdowns in situational perception. We introduce Fire360, a benchmark for evaluating perception and reasoning in safety-critical firefighting scenarios. The dataset includes 228 360 videos from professional training sessions under diverse conditions (e.g., low light, thermal distortion), annotated with action segments, object locations, and degradation metadata. Fire360 supports five tasks: Visual Question Answering, Temporal Action Captioning, Object Localization, Safety-Critical Reasoning, and Transformed Object Retrieval (TOR). TOR tests whether models can match pristine exemplars to fire-damaged counterparts in unpaired scenes, evaluating episodic memory under irreversible visual transformations. While human experts achieve 83.5% on TOR, models like GPT-4o lag significantly, exposing failures in reasoning under degradation. By releasing Fire360 and its evaluation suite, we aim to advance models that not only see, but also remember, reason, and act under uncertainty.
HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class
Large language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or involve formal proofs, often overlooking approximation-based problems ubiquitous in applied science and engineering. To fill this gap, we build on prior work and present $\textbf{HARDMath2}$, a dataset of 211 original problems covering the core topics in an introductory graduate applied math class, including boundary-layer analysis, WKB methods, asymptotic solutions of nonlinear partial differential equations, and the asymptotics of oscillatory integrals. This dataset was designed and verified by the students and instructors of a core graduate applied mathematics course at Harvard. We build the dataset through a novel collaborative environment that challenges students to write and refine difficult problems consistent with the class syllabus, peer-validate solutions, test different models, and automatically check LLM-generated solutions against their own answers and numerical ground truths. Evaluation results show that leading frontier models still struggle with many of the problems in the dataset, highlighting a gap in the mathematical reasoning skills of current LLMs. Importantly, students identified strategies to create increasingly difficult problems by interacting with the models and exploiting common failure modes. This back-and-forth with the models not only resulted in a richer and more challenging benchmark but also led to qualitative improvements in the students' understanding of the course material, which is increasingly important as we enter an age where state-of-the-art language models can solve many challenging problems across a wide domain of fields.
Get officially certified in Claude AI for just 19.99
When you purchase through links in our articles, we may earn a small commission. Get officially certified in Claude AI for just $19.99 A Claude AI Professional E-Degree is on sale for $19.99 (reg. AI skills are no longer a nice-to-have. A verifiable credential in one of the most popular AI models on the market is a real resume differentiator, and right now, you can get an e-degree in Claude for just $19.99 (reg. While plenty of people have dabbled with Claude, there's a big difference between "I've used it a few times" and actually knowing how to make it work for you.
Multimarginal flow matching with optimal transport potentials
Kansal, Raghav, Crair, David, Nguyen, Nghia, Pope, Scott, Parry, Bradley
Flow matching (FM) has emerged as a powerful framework for learning dynamic transport maps between two empirical distributions. However, less explored is the setting with intermediate observed marginals that can help constrain the flows between the endpoints. This "multimarginal" regime is central to modeling temporal evolution in dynamical systems in many scientific domains that can sample sequential distributions. We tackle this problem with a novel approach that leverages the connection between FM and dynamic optimal transport (OT), softly steering the flow towards the intermediate marginals through potential terms in the dynamic OT action. By extending the conditional FM learning target to incorporate these potentials, we derive an efficient, simulation-free algorithm for multimarginal FM that offers considerable flexibility in the spatiotemporal dynamics of the learned flows. We demonstrate state-of-the-art performance and training efficiency of OT-potential FM (OTP-FM) on diverse single-cell RNA sequencing, oceanographic, and meteorological datasets. Our code is available at https://github.com/Bexorg-Inc/OTP-FM.
Multicalibration Boosting: Theory, Convergence, and Transferability
Multicalibration extends classical calibration by requiring predictions to be unbiased over a rich collection of functions, encompassing both prediction slices and subpopulations. It has emerged as a powerful framework for fairness, robustness, and reliable prediction, yet the theoretical understanding of multicalibration boosting (MCBoost) remains fragmented and often relies on restrictive assumptions. In this work, we develop a unified and refined perspective on MCBoost that subsumes existing variants, including multiaccuracy, BatchGCP, and BatchMVP. We uncover several phenomena that provide new insights into its practical behavior: even highly accurate and flexible predictors can remain substantially miscalibrated; enforcing multicalibration introduces a calibration-risk trade-off; and early stopping plays a central role in controlling this trade-off. On the theoretical side, we establish a general framework for MCBoost under weaker and more realistic conditions. We show that the boosting iterates converge to a Bregman projection of the population-optimal predictor onto the cumulative span generated by the audit class, thereby explicitly characterizing the function space on which multicalibration is achieved. We further derive convergence rates under different smoothness assumptions, finite-sample guarantees, and principled stopping rules that ensure multicalibration at termination. Finally, we extend the theory of universal adaptability under covariate shift, providing more general transfer guarantees and clarifying when multicalibrated predictors generalize across domains. These results provide a more complete theoretical foundation and practical guidance for multicalibration boosting, positioning it as both a unifying framework and a reliable post-processing approach for modern predictive models.
Concomitant DAG Learning: On the Roles of Noise Adaptivity, Sparsity, and Non-negativity
Mateos, Gonzalo, Rey, Samuel, Ajorlou, Hamed, Tepper, Mariano
Directed acyclic graphs (DAGs) constitute a central modeling tool to enable principled reasoning about cause-effect interactions in complex systems. However, since the causal structure underlying a group of variables is often unknown and interventions may be infeasible or ethically challenging to implement, there is a need to address the task of inferring DAGs from observational data. However, most classical structure identification approaches face two key obstacles: the combinatorial challenge of enforcing acyclicity, which severely limits scalability, and identifiability challenges arising from latent confounding or heterogeneous noise. This tutorial offers an overview of recent signal processing and optimization advances that address these issues by recasting DAG structure learning as a continuous, score-based estimation problem over adjacency matrices. We begin with a didactic introduction to structural equation models and the formulation of causal graph recovery, followed by a historical survey of score-based methods ranging from early combinatorial search schemes and greedy heuristics to modern continuous frameworks that leverage smooth characterizations of acyclicity. Building on this foundation, we describe concomitant DAG estimation methods that jointly infer sparse causal structure and exogenous noise levels, improving robustness under heteroscedasticity and distribution shifts by rendering the estimator noise adaptive. All in all, the tutorial introduces readers to challenges and opportunities for signal processing research at the crossroads of causal inference, high-dimensional statistics, and scalable graph learning, while outlining emerging directions including online, nonlinear, and neural causal discovery.
Get the newest Microsoft dev tools plus 15 coding courses -- only 50
When you purchase through links in our articles, we may earn a small commission. TL;DR: The Microsoft Visual Studio Professional 2026 bundle includes 15 coding courses and is on sale for $49.97 (regularly $1,999.99) This is the kind of tech purchase that tends to pay for itself pretty quickly. The Microsoft Visual Studio Professional 2026 bundle pairs one of the most widely used development environments in the industry with a full library of coding courses -- all for a single one-time payment. Visual Studio Professional 2026 has been a staple for professional developers for years, and the 2026 version pushes productivity even further.
The Zuckerbergs Are Hiring a Lifeguard but Calling It a 'Beach Water Person'
The Zuckerbergs Are Hiring a Lifeguard but Calling It a'Beach Water Person' The job, which is associated with the Zuckerberg family office, is located in Kauai, Hawaii, where the Meta CEO owns a massive compound. Meta CEO Mark Zuckerberg and his wife Priscilla Chan are hiring a seasonal, on-call "Beach Water Person" based in Kauai, Hawaii, where the family owns a sprawling compound, according to a new job listing on Greenhouse associated with West 10, the Zuckerberg family office. This is an interesting choice for a job title, because according to the job description, the primary duties of this "Beach Water Person" include serving as a "Beach Lifeguard," and "Pool Lifeguard." The job listing names a few additional duties related to water activities, such as instructing "stand-up paddleboarding (SUP), canoe paddling, snorkeling, and other ocean-based activities." These, however, come after the water safety duties in the job description.
An ICE Firearms Trainer Was Involved in At Least 4 Deadly Shootings
David Norman, a former Phoenix police officer who's described himself as "a fucking savage," now runs a company that provided training to Homeland Security's Special Response Teams. The owner of a company that trained paramilitary Immigration and Customs Enforcement agents testified that he was involved in at least four lethal shootings, according to a 2021 deposition related to a lawsuit reviewed by WIRED. David S. Norman, the founder and proprietor of law enforcement training firm TruKinetics LLC, served as a Phoenix Police officer from the late 1990s until his retirement in 2020. Prior to founding TruKinetics the same year, according to records reviewed by WIRED, Norman was involved in six shootings while on duty that left four people dead and two more wounded. In every instance, the Phoenix Police Department said Norman fired on an armed suspect and exchanged volleys of gunfire in at least two of the shootings. Based in Gilbert, Arizona, TruKinetics offers training on small-team tactics, hostage rescues, close-quarters combat, building searches, night-vision firearms proficiency, pistol and rifle courses, "vehicle interdiction," breaching with explosives, and sniper tactics, according to the company's website.