Goto

Collaborating Authors

 Deep Learning


Optimization, Generalization and Differential Privacy Bounds for Gradient Descent on Kolmogorov-Arnold Networks

arXiv.org Machine Learning

Kolmogorov--Arnold Networks (KANs) have recently emerged as a structured alternative to standard MLPs, yet a principled theory for their training dynamics, generalization, and privacy properties remains limited. In this paper, we analyze gradient descent (GD) for training two-layer KANs and derive general bounds that characterize their training dynamics, generalization, and utility under differential privacy (DP). As a concrete instantiation, we specialize our analysis to logistic loss under an NTK-separable assumption, where we show that polylogarithmic network width suffices for GD to achieve an optimization rate of order $1/T$ and a generalization rate of order $1/n$, with $T$ denoting the number of GD iterations and $n$ the sample size. In the private setting, we characterize the noise required for $(ε,δ)$-DP and obtain a utility bound of order $\sqrt{d}/(nε)$ (with $d$ the input dimension), matching the classical lower bound for general convex Lipschitz problems. Our results imply that polylogarithmic width is not only sufficient but also necessary under differential privacy, revealing a qualitative gap between non-private (sufficiency only) and private (necessity also emerges) training regimes. Experiments further illustrate how these theoretical insights can guide practical choices, including network width selection and early stopping.


Provable Target Sample Complexity Improvements as Pre-Trained Models Scale

arXiv.org Machine Learning

Pre-trained models have become indispensable for efficiently building models across a broad spectrum of downstream tasks. The advantages of pre-trained models have been highlighted by empirical studies on scaling laws, which demonstrate that larger pre-trained models can significantly reduce the sample complexity of downstream learning. However, existing theoretical investigations of pre-trained models lack the capability to explain this phenomenon. In this paper, we provide a theoretical investigation by introducing a novel framework, caulking, inspired by parameter-efficient fine-tuning (PEFT) methods such as adapter-based fine-tuning, low-rank adaptation, and partial fine-tuning. Our analysis establishes that improved pre-trained models provably decrease the sample complexity of downstream tasks, thereby offering theoretical justification for the empirically observed scaling laws relating pre-trained model size to downstream performance, a relationship not covered by existing results.


Group Contrastive Learning for Weakly Paired Multimodal Data

arXiv.org Machine Learning

We present GROOVE, a semi-supervised multi-modal representation learning approach for high-content perturbation data where samples across modalities are weakly paired through shared perturbation labels but lack direct correspondence. Our primary contribution is GroupCLIP, a novel group-level contrastive loss that bridges the gap between CLIP for paired cross-modal data and SupCon for uni-modal supervised contrastive learning, addressing a fundamental gap in contrastive learning for weakly-paired settings. We integrate GroupCLIP with an on-the-fly backtranslating autoencoder framework to encourage cross-modally entangled representations while maintaining group-level coherence within a shared latent space. Critically, we introduce a comprehensive combinatorial evaluation framework that systematically assesses representation learners across multiple optimal transport aligners, addressing key limitations in existing evaluation strategies. This framework includes novel simulations that systematically vary shared versus modality-specific perturbation effects enabling principled assessment of method robustness. Our combinatorial benchmarking reveals that there is not yet an aligner that uniformly dominates across settings or modality pairs. Across simulations and two real single-cell genetic perturbation datasets, GROOVE performs on par with or outperforms existing approaches for downstream cross-modal matching and imputation tasks. Our ablation studies demonstrate that GroupCLIP is the key component driving performance gains. These results highlight the importance of leveraging group-level constraints for effective multi-modal representation learning in scenarios where only weak pairing is available.


A New AI Math Startup Just Cracked 4 Previously Unsolved Problems

WIRED

Axiom says its AI found solutions to several long-standing math problems, a sign of the technology's steadily advancing reasoning capabilities. Five years ago, mathematicians Dawei Chen and Quentin Gendron were trying to untangle a difficult area of algebraic geometry involving differentials, elements of calculus used to measure distance along curved surfaces . While working on one theorem, they ran into an unexpected roadblock: Their argument depended on a strange formula from number theory, but they were unable to solve or justify it. In the end, Chen and Gendron wrote a paper presenting their idea as a conjecture, rather than a theorem. Chen recently spent hours prompting ChatGPT in the hopes of getting the AI to come up with a solution to the still unsolved problem, but it wasn't working.


Microsoft Copilot claims it can set reminders. My phone never buzzed

PCWorld

PCWorld tested Microsoft Copilot's new reminder feature for Android and iOS phones, which allows setting reminders from a PC similar to old Cortana functions. The feature proved unreliable during testing, with reminders failing to trigger notifications on devices, raising concerns about Copilot's overall utility. With SimilarWeb reporting only 1 percent usage figures for Copilot, this unreliability could further impact user trust and adoption rates. Microsoft has quietly added reminders to Copilot. Well, at least Copilot seems to think so.


Anthropic says it won't bring ads to Claude, unlike rival ChatGPT

Engadget

Anthropic says it won't bring ads to Claude, unlike rival ChatGPT The company said that integrating advertising would'work against' the core principles of the chatbot. Anthropic has announced that its chatbot Claude . This is in direct contrast to rival company OpenAI, which recently for many users. The company says that including ads in conversations with Claude would be incompatible with the chatbot becoming a genuinely helpful assistant for work and for deep thinking. The reasoning here is rather simple.


Amazon's new AI Alexa isn't free anymore

PCWorld

PCWorld reports that Amazon has ended free access to its advanced Alexa+ AI service, now charging non-Prime members $19.99 monthly for full features. Prime subscribers receive Alexa+ at no additional cost, while the original classic Alexa remains free for all users regardless of membership status. Alexa+ offers ChatGPT-style conversations and advanced agentic AI capabilities, officially exiting its early access phase with a limited free web-based text option available. The days of free Alexa+ for everyone just ended, with Amazon announcing today that it will start charging non-Prime members who want to use the AI-supercharged voice assistant on their Echo devices. Starting now, full Alexa+ access will cost $19.99 a month for those without Prime, while Prime members will get Alexa+ as a free benefit with their subscriptions. Amazon also announced a new free tier of Alexa+ that lets you text chat with the assistant over a web browser.


The Download: the future of nuclear power plants, and social media-fueled AI hype

MIT Technology Review

AI is driving unprecedented investment for massive data centers and an energy supply that can support its huge computational appetite. One potential source of electricity for these facilities is next-generation nuclear power plants, which could be cheaper to construct and safer to operate than their predecessors. We recently held a subscriber-exclusive Roundtables discussion on hyperscale AI data centers and next-gen nuclear --two featured technologies on the MIT Technology Review 10 Breakthrough Technologies of 2026 list . You can watch the conversation back here, and don't forget to subscribe to make sure you catch future discussions as they happen. Demis Hassabis, CEO of Google DeepMind, summed it up in three words: "This is embarrassing." Hassabis was replying on X to an overexcited post by Sébastien Bubeck, a research scientist at the rival firm OpenAI, announcing that two mathematicians had used OpenAI's latest large language model, GPT-5, to find solutions to 10 unsolved problems in mathematics.


AI Bots Are Now a Signifigant Source of Web Traffic

WIRED

New data shows AI bots pushing deeper into the web, prompting publishers to roll out more aggressive defenses. The viral virtual assistant OpenClaw--formerly known as Moltbot, and before that Clawdbot--is a symbol of a broader revolution underway that could fundamentally alter how the internet functions. Instead of a place primarily inhabited by humans, the web may very soon be dominated by autonomous AI bots. A new report measuring bot activity on the web, as well as related data shared with WIRED by the internet infrastructure company Akamai, shows that AI bots already account for a meaningful share of web traffic. The findings also shed light on an increasingly sophisticated arms race unfolding as bots deploy clever tactics to bypass website defenses meant to keep them out.


HHS Is Making an AI Tool to Create Hypotheses About Vaccine Injury Claims

WIRED

Experts worry Robert F. Kennedy Jr.'s Health Department will use an internal AI tool to analyze vaccine injury claims in a way that furthers his anti-vaccine agenda. The US Department of Health and Human Services is developing a generative artificial intelligence tool to find patterns across data reported to a national vaccine monitoring database and to generate hypotheses on the negative effects of vaccines, according to an inventory released last week of all use cases the agency had for AI in 2025. The tool has not yet been deployed, according to the HHS document, and an AI inventory report from the previous year shows that it has been in development since late 2023. But experts worry that the predictions it generates could be used by Health and Human Services secretary Robert F. Kennedy Jr. to further his anti-vaccine agenda. A long-standing vaccine critic, Kenedy has upended the childhood vaccination schedule in his year in office, removing several shots from a list of recommended immunizations for all children, including those for Covid-19, influenza, hepatitis A and B, meningococcal disease, rotavirus, and respiratory syncytial virus, or RSV.