Deep Learning
Monitoring the calibration of probability forecasts with an application to concept drift detection involving image classification
Franck, Christopher T., Driscoll, Anne R., Szajnfarber, Zoe, Woodall, William H.
Machine learning approaches for image classification have led to impressive advances in that field. For example, convolutional neural networks are able to achieve remarkable image classification accuracy across a wide range of applications in industry, defense, and other areas. While these machine learning models boast impressive accuracy, a related concern is how to assess and maintain calibration in the predictions these models make. A classification model is said to be well calibrated if its predicted probabilities correspond with the rates events actually occur. While there are many available methods to assess machine learning calibration and recalibrate faulty predictions, less effort has been spent on developing approaches that continually monitor predictive models for potential loss of calibration as time passes. We propose a cumulative sum-based approach with dynamic limits that enable detection of miscalibration in both traditional process monitoring and concept drift applications. This enables early detection of operational context changes that impact image classification performance in the field. The proposed chart can be used broadly in any situation where the user needs to monitor probability predictions over time for potential lapses in calibration. Importantly, our method operates on probability predictions and event outcomes and does not require under-the-hood access to the machine learning model.
Generative Bayesian Optimization: Generative Models as Acquisition Functions
Oliveira, Rafael, Steinberg, Daniel M., Bonilla, Edwin V.
We present a general strategy for turning generative models into candidate solution samplers for batch Bayesian optimization (BO). The use of generative models for BO enables large batch scaling as generative sampling, optimization of non-continuous design spaces, and high-dimensional and combinatorial design. Inspired by the success of direct preference optimization (DPO), we show that one can train a generative model with noisy, simple utility values directly computed from observations to then form proposal distributions whose densities are proportional to the expected utility, i.e., BO's acquisition function values. Furthermore, this approach is generalizable beyond preference-based feedback to general types of reward signals and loss functions. This perspective avoids the construction of surrogate (regression or classification) models, common in previous methods that have used generative models for black-box optimization. Theoretically, we show that the generative models within the BO process approximately follow a sequence of distributions which asymptotically concentrate at the global optima under certain conditions. We also demonstrate this effect through experiments on challenging optimization problems involving large batches in high dimensions.
From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning
Oh, Junsoo, Song, Jerry, Yun, Chulhee
Weak-to-strong generalization refers to the phenomenon where a stronger model trained under supervision from a weaker one can outperform its teacher. While prior studies aim to explain this effect, most theoretical insights are limited to abstract frameworks or linear/random feature models. In this paper, we provide a formal analysis of weak-to-strong generalization from a linear CNN (weak) to a two-layer ReLU CNN (strong). We consider structured data composed of label-dependent signals of varying difficulty and label-independent noise, and analyze gradient descent dynamics when the strong model is trained on data labeled by the pretrained weak model. Our analysis identifies two regimes -- data-scarce and data-abundant -- based on the signal-to-noise characteristics of the dataset, and reveals distinct mechanisms of weak-to-strong generalization. In the data-scarce regime, generalization occurs via benign overfitting or fails via harmful overfitting, depending on the amount of data, and we characterize the transition boundary. In the data-abundant regime, generalization emerges in the early phase through label correction, but we observe that overtraining can subsequently degrade performance.
The Sign Estimator: LLM Alignment in the Face of Choice Heterogeneity
Aouad, Ali, Gadarri, Aymane El, Farias, Vivek F.
Traditional LLM alignment methods are vulnerable to heterogeneity in human preferences. Fitting a naรฏve probabilistic model to pairwise comparison data (say over prompt-completion pairs) yields an inconsistent estimate of the population-average utility -a canonical measure of social welfare. We propose a new method, dubbed the sign estimator, that provides a simple, provably consistent, and efficient estimator by replacing cross-entropy with binary classification loss in the aggregation step. This simple modification recovers consistent ordinal alignment under mild assumptions and achieves the first polynomial finite-sample error bounds in this setting. In realistic simulations of LLM alignment using digital twins, the sign estimator substantially reduces preference distortion over a panel of simulated personas, cutting (angular) estimation error by nearly 35% and decreasing disagreement with true population preferences from 12% to 8% compared to standard RLHF. Our method also compares favorably to panel data heuristics that explicitly model user heterogeneity and require tracking individual-level preference data-all while maintaining the implementation simplicity of existing LLM alignment pipelines.
CANDI: Hybrid Discrete-Continuous Diffusion Models
Pynadath, Patrick, Shi, Jiaxin, Zhang, Ruqi
While continuous diffusion has shown remarkable success in continuous domains such as image generation, its direct application to discrete data has underperformed compared to purely discrete formulations. This gap is counterintuitive, given that continuous diffusion learns score functions that enable joint evolution across multiple positions. To understand this gap, we introduce token identifiability as an analytical framework for understanding how Gaussian noise corrupts discrete data through two mechanisms: discrete identity corruption and continuous rank degradation. We reveal that these mechanisms scale differently with vocabulary size, creating a temporal dissonance: at noise levels where discrete corruption preserves enough structure for conditional learning, continuous denoising is trivial; at noise levels where continuous denoising is meaningful, discrete corruption destroys nearly all conditional structure. To solve this, we propose CANDI (Continuous ANd DIscrete diffusion), a hybrid framework that decouples discrete and continuous corruption, enabling simultaneous learning of both conditional structure and continuous geometry. We empirically validate the temporal dissonance phenomenon and demonstrate that CANDI successfully avoids it. This unlocks the benefits of continuous diffusion for discrete spaces: on controlled generation, CANDI enables classifier-based guidance with off-the-shelf classifiers through simple gradient addition; on text generation, CANDI outperforms masked diffusion at low NFE, demonstrating the value of learning continuous gradients for discrete spaces. We include the code on the project page available here: https://patrickpynadath1.github.io/candi-lander
Microsoft reports strong earnings as Azure hit by major outage
Microsoft's CEO, Satya Nadella, speaks at the company's annual developer conference in Seattle, Washington. Microsoft's CEO, Satya Nadella, speaks at the company's annual developer conference in Seattle, Washington. Tech giant reports earnings of $3.72 per share day after deal with OpenAI pushed value of company to more than $4tn Microsoft blew off concerns of overspending on AI on Wednesday, reporting elevated earnings even as it faced an outage of its cloud computing service, Azure, and its office software suite, 365. The strong earnings report comes a day after a deal with OpenAI pushed the value of the tech giant to more than $4tn. After its Xbox and investor relations pages went down, the company issued a statement that said: "We are working to address an issue affecting Azure Front Door that is impacting the availability of some services."
AI Agents Are Terrible Freelance Workers
Human-level AI is still some ways off. Even the best artificial intelligence agents are fairly hopeless at online freelance work, according to an experiment that challenges the idea of AI replacing office workers en masse. The Remote Labor Index, a new benchmark developed by researchers at data annotation company Scale AI and the Center for AI Safety (CAIS), a nonprofit, measures the ability of frontier AI models to automate economically valuable work. The researchers gave several leading AI agents a range of simulated freelance work and found that even the best could perform less than 3 percent of the work, earning $1,810 out of a possible $143,991. The researchers looked at several tools and found the most capable to be Manus from a Chinese startup of the same name, followed by Grok from xAI, Claude from Anthropic, ChatGPT from OpenAI, and Gemini from Google.
ChatGPT teams up with PayPal to make it easier for you to buy stuff in chat
When you purchase through links in our articles, we may earn a small commission. Users will soon be able to use PayPal to pay for product recommendations made by OpenAI's ChatGPT. PayPal recently signed a contract with OpenAI to integrate the digital wallet into ChatGPT, reports CNBC . This will allow users to easily pay for the products they discover via the AI tool. The agreement allows PayPal users to make payments via ChatGPT merchants to list and sell their goods in ChatGPT.
Building a high performance data and AI organization (2nd edition)
What it takes to deliver on data and AI strategy. Four years is a lifetime when it comes to artificial intelligence. Since the first edition of this study was published in 2021, AI's capabilities have been advancing at speed, and the advances have not slowed since generative AI's breakthrough. For example, multimodality-- the ability to process information not only as text but also as audio, video, and other unstructured formats--is becoming a common feature of AI models. AI's capacity to reason and act autonomously has also grown, and organizations are now starting to work with AI agents that can do just that. Amid all the change, there remains a constant: the quality of an AI model's outputs is only ever as good as the data that feeds it.
The Download: Boosting AI's memory, and data centers' unhappy neighbors
DeepSeek may have found a new way to improve AI's ability to remember An AI model released by Chinese AI company DeepSeek uses new techniques that could significantly improve AI's ability to "remember." The optical character recognition model works by extracting text from an image and turning it into machine-readable words. This is the same technology that powers scanner apps, translation of text in photos, and many accessibility tools. Researchers say the model's main innovation lies in how it processes information--specifically, how it stores and retrieves data. Improving how AI models "remember" could reduce how much computing power they need to run, thus mitigating AI's large (and growing) carbon footprint. The AI Hype Index: Data centers' neighbors are pivoting to power blackouts That's why we've created the AI Hype Index--a simple, at-a-glance summary of everything you need to know about the state of the industry.