Deep Learning
Inside OpenAI's big play for science
An exclusive conversation with Kevin Weil, head of OpenAI for Science, a new in-house team that wants to make scientists more productive. In the three years since ChatGPT's explosive debut, OpenAI's technology has upended a remarkable range of everyday activities at home, at work, in schools--anywhere people have a browser open or a phone out, which is everywhere. Now OpenAI is making an explicit play for scientists. In October, the firm announced that it had launched a whole new team, called OpenAI for Science, dedicated to exploring how its large language models could help scientists and tweaking its tools to support them. The last couple of months have seen a slew of social media posts and academic publications in which mathematicians, physicists, biologists, and others have described how LLMs (and OpenAI's GPT-5 in particular) have helped them make a discovery or nudged them toward a solution they might otherwise have missed. In part, OpenAI for Science was set up to engage with this community.
How to generate AI images using ChatGPT
Apple could unveil Gemini-powered Siri in Feb. A good prompt goes a long way. ChatGPT is available on both iOS and Android. Since March 2025, ChatGPT has been capable of generating images. Following a period where it briefly wasn't available to free users, you now don't even pay for one of OpenAI's subscriptions to use this feature.
Why chatbots are starting to check your age
Confirming which users are kids is politically fraught and a technical nightmare. Here's what moves from OpenAI and the FTC tell us. How do tech companies check if their users are kids? This question has taken on new urgency recently thanks to growing concern about the dangers that can arise when children talk to AI chatbots. For years Big Tech asked for birthdays (that one could make up) to avoid violating child privacy laws, but they weren't required to moderate content accordingly. Two developments over the last week show how quickly things are changing in the US and how this issue is becoming a new battleground, even among parents and child-safety advocates.
ChatGPT is now indexing Grok's AI slop
PCWorld reports that ChatGPT 5.2 is now indexing Grokipedia, xAI's AI-generated encyclopedia known for inaccuracies and conspiracy theories. This creates a concerning feedback loop where AI-generated misinformation spreads between major language models, potentially overwriting established knowledge. The integration poses significant risks to information integrity as biased or false content from one AI system influences another's responses. More and more of the web is filling up with LLM-generated text, images, and even videos and music . It's an even bigger problem than it seems because the "AI" systems that have scoured the web to generate their large language models are now .
UK maker of AI avatars nearly doubles valuation to 4bn after funding round
A British AI startup that makes realistic video avatars has almost doubled its valuation to $4bn (ยฃ3bn), in a boost for the UK technology sector. Synthesia was valued at $2.1bn last year and moved into new offices in central London, marking the moment with a ceremony attended by the Sadiq Khan, the city's mayor, and Peter Kyle, then technology secretary. On Monday, it announced its latest funding round, led by an existing investor, Google Ventures, had raised $200m and valued the British company at $4bn. Google Ventures is the search firm's venture capital arm. Synthesia uses human actors to generate digital avatars of people and also offers employers the ability to create replicas of their staff.
FedSGM: A Unified Framework for Constraint Aware, Bidirectionally Compressed, Multi-Step Federated Optimization
Upadhyay, Antesh, Moon, Sang Bin, Hashemi, Abolfazl
We introduce FedSGM, a unified framework for federated constrained optimization that addresses four major challenges in federated learning (FL): functional constraints, communication bottlenecks, local updates, and partial client participation. Building on the switching gradient method, FedSGM provides projection-free, primal-only updates, avoiding expensive dual-variable tuning or inner solvers. To handle communication limits, FedSGM incorporates bi-directional error feedback, correcting the bias introduced by compression while explicitly understanding the interaction between compression noise and multi-step local updates. We derive convergence guarantees showing that the averaged iterate achieves the canonical $\boldsymbol{\mathcal{O}}(1/\sqrt{T})$ rate, with additional high-probability bounds that decouple optimization progress from sampling noise due to partial participation. Additionally, we introduce a soft switching version of FedSGM to stabilize updates near the feasibility boundary. To our knowledge, FedSGM is the first framework to unify functional constraints, compression, multiple local updates, and partial client participation, establishing a theoretically grounded foundation for constrained federated learning. Finally, we validate the theoretical guarantees of FedSGM via experimentation on Neyman-Pearson classification and constrained Markov decision process (CMDP) tasks.
Multigrade Neural Network Approximation
Zhang, Shijun, Shen, Zuowei, Xu, Yuesheng
We study multigrade deep learning (MGDL) as a principled framework for structured error refinement in deep neural networks. While the approximation power of neural networks is now relatively well understood, training very deep architectures remains challenging due to highly non-convex and often ill-conditioned optimization landscapes. In contrast, for relatively shallow networks, most notably one-hidden-layer $\texttt{ReLU}$ models, training admits convex reformulations with global guarantees, motivating learning paradigms that improve stability while scaling to depth. MGDL builds upon this insight by training deep networks grade by grade: previously learned grades are frozen, and each new residual block is trained solely to reduce the remaining approximation error, yielding an interpretable and stable hierarchical refinement process. We develop an operator-theoretic foundation for MGDL and prove that, for any continuous target function, there exists a fixed-width multigrade $\texttt{ReLU}$ scheme whose residuals decrease strictly across grades and converge uniformly to zero. To the best of our knowledge, this work provides the first rigorous theoretical guarantee that grade-wise training yields provable vanishing approximation error in deep networks. Numerical experiments further illustrate the theoretical results.
Long-Term Probabilistic Forecast of Vegetation Conditions Using Climate Attributes in the Four Corners Region
McPhillips, Erika, Lee, Hyeongseong, Xie, Xiangyu, Baylis, Kathy, Funk, Chris, Gu, Mengyang
Weather conditions can drastically alter the state of crops and rangelands, and in turn, impact the incomes and food security of individuals worldwide. Satellite-based remote sensing offers an effective way to monitor vegetation and climate variables on regional and global scales. The annual peak Normalized Difference Vegetation Index (NDVI), derived from satellite observations, is closely associated with crop development, rangeland biomass, and vegetation growth. Although various machine learning methods have been developed to forecast NDVI over short time ranges, such as one-month-ahead predictions, long-term forecasting approaches, such as one-year-ahead predictions of vegetation conditions, are not yet available. To fill this gap, we develop a two-phase machine learning model to forecast the one-year-ahead peak NDVI over high-resolution grids, using the Four Corners region of the Southwestern United States as a testbed. In phase one, we identify informative climate attributes, including precipitation and maximum vapor pressure deficit, and develop the generalized parallel Gaussian process that captures the relationship between climate attributes and NDVI. In phase two, we forecast these climate attributes using historical data at least one year before the NDVI prediction month, which then serve as inputs to forecast the peak NDVI at each spatial grid. We developed open-source tools that outperform alternative methods for both gross NDVI and grid-based NDVI one-year forecasts, providing information that can help farmers and ranchers make actionable plans a year in advance.
Towards Latent Diffusion Suitable For Text
Midavaine, Nesta, Naesseth, Christian A., Bartosh, Grigory
Language diffusion models aim to improve sampling speed and coherence over autoregressive LLMs. We introduce Neural Flow Diffusion Models for language generation, an extension of NFDM that enables the straightforward application of continuous diffusion models to discrete state spaces. NFDM learns a multivariate forward process from the data, ensuring that the forward process and generative trajectory are a good fit for language modeling. Our model substantially reduces the likelihood gap with autoregressive models of the same size, while achieving sample quality comparable to that of previous latent diffusion models.
Apple reportedly plans to reveal its Gemini-powered Siri in February
Bloomberg reports that Apple will show off demonstrations of the revamped Siri in the second half of February. A new and improved Siri may finally make an appearance, but this time, it could be with a Google Gemini glow up. According to's Mark Gurman, Apple wants to announce a new Siri in the second half of February that will show off the results of its recently announced partnership with Google and offer demonstrations of the Gemini-powered capabilities. After this reveal, Gurman reported that the new Siri will make its way to iOS 26.4, Apple has been meaning to launch its next-gen Siri ever since its announcement at WWDC 2024, but now we know that this Gemini-powered Siri will behave more like an AI chatbot, similar to OpenAI's ChatGPT, thanks to another report from last week.