Deep Learning
The Download: the LLM will see you now, and a new fusion power deal
Patients at a small number of clinics in Southern California run by the medical startup Akido Labs are spending relatively little time, or even no time at all, with their doctors. Instead, they see a medical assistant, who can lend a sympathetic ear but has limited clinical training. The job of formulating diagnoses and concocting a treatment plan is done by an LLM-based system called ScopeAI that transcribes and analyzes the dialogue between patient and assistant. A doctor then approves, or corrects, the AI system's recommendations. According to Akido's CEO, this approach allows doctors to see four to five times as many patients as they could previously. But experts aren't convinced that displacing so much of the cognitive work of medicine onto AI is the right way to remedy the doctor shortage.
If A.I. Can Diagnose Patients, What Are Doctors For?
If A.I. Can Diagnose Patients, What Are Doctors For? Large language models are transforming medicine--but the technology comes with side effects. "I'm worried these tools will erode my ability to make an independent diagnosis," a medical student said. In 2017, Matthew Williams, a thirtysomething software engineer with an athletic build and a bald head, went for a long bike ride in the hills of San Francisco. Afterward, at dinner with some friends, he ordered a hamburger, fries, and a milkshake. Midway through the meal, he felt so full that he had to ask someone to drive him home. That night, Williams awoke with a sharp pain in his abdomen that he worried was appendicitis. He went to a nearby emergency clinic, where doctors told him that he was probably constipated. They gave him some laxatives and sent him on his way. A few hours later, Williams's pain intensified. He vomited and felt as though his stomach might burst. A friend took him to a hospital, where a CT scan revealed cecal volvulus--a medical emergency in which part of the intestine twists in on itself, cutting off the digestive tract. The previous medical team had missed the condition, and may even have exacerbated it by giving him laxatives. Williams was rushed to the operating room, where surgeons removed about six feet of his intestines. After recovering from surgery, Williams began to experience severe diarrhea almost every time he ate. Doctors told him that his bowel just needed time to heal. "It got to the point where I couldn't go out, because I would constantly eat something that would make me sick," he said.
This medical startup uses LLMs to run appointments and make diagnoses
"Our focus is really on what we can do to pull the doctor out of the visit," says Akido's CTO. Imagine this: You've been feeling unwell, so you call up your doctor's office to make an appointment. At the appointment, you aren't rushed through describing your health concerns; instead, you have a full half hour to share your symptoms and worries and the exhaustive details of your health history with someone who listens attentively and asks thoughtful follow-up questions. You leave with a diagnosis, a treatment plan, and the sense that, for once, you've been able to discuss your health with the care that it merits. AI companies have stopped warning you that their chatbots aren't doctors Once cautious, OpenAI, Grok, and others will now dive into giving unverified medical advice with virtually no disclaimers. You might not have spoken to a doctor, or other licensed medical practitioner, at all.
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
Li, Pengteng, Song, Pinhao, Li, Wuyang, Guo, Weiyu, Yao, Huizai, Xu, Yijie, Liu, Dugang, Xiong, Hui
We introduce SEE&TREK, the first training-free prompting framework tailored to enhance the spatial understanding of Multimodal Large Language Models (MLLMS) under vision-only constraints. While prior efforts have incorporated modalities like depth or point clouds to improve spatial reasoning, purely visualspatial understanding remains underexplored. SEE&TREK addresses this gap by focusing on two core principles: increasing visual diversity and motion reconstruction. For visual diversity, we conduct Maximum Semantic Richness Sampling, which employs an off-the-shell perception model to extract semantically rich keyframes that capture scene structure. For motion reconstruction, we simulate visual trajectories and encode relative spatial positions into keyframes to preserve both spatial relations and temporal coherence. Our method is training&GPU-free, requiring only a single forward pass, and can be seamlessly integrated into existing MLLM'S. Extensive experiments on the VSI-B ENCH and STI-B ENCH show that S EE &T REK consistently boosts various MLLM S performance across diverse spatial reasoning tasks with the most +3.5% improvement, offering a promising path toward stronger spatial intelligence.
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
Rajaee, Sara, Choenni, Rochelle, Shutova, Ekaterina, Monz, Christof
While the reasoning abilities of large language models (LLMs) continue to advance, it remains unclear how such ability varies across languages in multilingual LLMs and whether different languages produce reasoning paths that complement each other. To investigate this question, we train a reward model to rank generated responses for a given question across languages. Our results show that our cross-lingual reward model substantially improves mathematical reasoning performance compared to using reward modeling within a single language, benefiting even high-resource languages. While English often exhibits the highest performance in multilingual models, we find that cross-lingual sampling particularly benefits English under low sampling budgets. Our findings reveal new opportunities to improve multilingual reasoning by leveraging the complementary strengths of diverse languages.
RAVE: Retrieval and Scoring Aware Verifiable Claim Detection
ABSTRACT The rapid spread of misinformation on social media underscores the need for scalable fact-checking tools. A key step is claim detection, which identifies statements that can be objectively verified. Prior approaches often rely on linguistic cues or claim check-worthiness, but these struggle with vague political discourse and diverse formats such as tweets. We present RA VE (Retrieval and Scoring A ware V erifiable Claim Detection), a framework that combines evidence retrieval with structured signals of relevance and source credibility. Experiments on CT22-test and PoliClaim-test show that RA VE consistently outperforms text-only and retrieval-based baselines in both accuracy and F1.
Quantum Generative Adversarial Autoencoders: Learning latent representations for quantum data generation
Raj, Naipunnya, Sangle, Rajiv, Singh, Avinash, Sabapathy, Krishna Kumar
Over the past decade, machine learning has undergone transformative advancements, primarily fueled by the development of sophisticated deep learning architectures and training methodologies. In parallel, Quantum Machine Learning (QML) has emerged as a field dedicated to exploring how quantum algorithms and quantum computing platforms can be utilized to process, model, and extract meaningful insights from data [9, 14, 65], and also generate new data [26, 59]. While efforts in QML primarily focused on leveraging quantum computing to accelerate classical machine learning tasks [19, 34], a significant and increasingly important direction involves the development of quantum models that operate directly on quantum data [7, 9, 41]. These models, tailored specifically to quantum data, are essential for realizing the full potential of quantum technologies, enabling applications in quantum information processing that are intractable with classical methods [25]. A notable model within QML for handling quantum data is the Quantum Autoencoder (QAE), which draws inspiration from its classical counterpart, the Autoen-coder (AE) [5, 58]. QAE has been applied to demonstrate how quantum circuits can be trained to compress quantum states, with applications to quantum simulation and quantum information [13, 29, 42, 44, 57]. Further developments extend these architectures to the denoising of entangled quantum states under realistic noise models [1, 10, 62, 63], along with proposals for error mitigation strategies tailored to Noisy Intermediate-Scale Quantum (NISQ) devices [46, 66]. Practical realizations of QAE in quantum hardware, such as nitrogen-vacancy centers, demonstrated robust compression and the preservation of entanglement, while significantly lengthening the coherence times of Bell states [67]. These two authors contributed equally.
The Alignment Bottleneck
Large language models improve with scale, yet feedback-based alignment still exhibits systematic deviations from intended behavior. Motivated by bounded rationality in economics and cognitive science, we view judgment as resource-limited and feedback as a constrained channel. On this basis, we model the loop as a two-stage cascade $U \to H \to Y$ given $S$, with cognitive capacity $C_{\text{cog}|S}$ and average total capacity $\bar{C}_{\text{tot}|S}$. Our main result is a capacity-coupled Alignment Performance Interval. It pairs a data size-independent Fano lower bound proved on a separable codebook mixture with a PAC-Bayes upper bound whose KL term is controlled by the same channel via $m \, \bar{C}_{\text{tot}|S}$. The PAC-Bayes bound becomes an upper bound on the same true risk when the canonical observable loss is used and the dataset is drawn from the same mixture. Under these matched conditions, both limits are governed by a single capacity. Consequences include that, with value complexity and capacity fixed, adding labels alone cannot cross the bound; attaining lower risk on more complex targets requires capacity that grows with $\log M$; and once useful signal saturates capacity, further optimization tends to fit channel regularities, consistent with reports of sycophancy and reward hacking. The analysis views alignment as interface engineering: measure and allocate limited capacity, manage task complexity, and decide where information is spent.
Generalization and Optimization of SGD with Lookahead
The Lookahead optimizer enhances deep learning models by employing a dual-weight update mechanism, which has been shown to improve the performance of underlying optimizers such as SGD. However, most theoretical studies focus on its convergence on training data, leaving its generalization capabilities less understood. Existing generalization analyses are often limited by restrictive assumptions, such as requiring the loss function to be globally Lipschitz continuous, and their bounds do not fully capture the relationship between optimization and generalization. In this paper, we address these issues by conducting a rigorous stability and generalization analysis of the Lookahead optimizer with minibatch SGD. We leverage on-average model stability to derive generalization bounds for both convex and strongly convex problems without the restrictive Lipschitzness assumption. Our analysis demonstrates a linear speedup with respect to the batch size in the convex setting.
Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
Lin, Zinan, Liu, Enshu, Ning, Xuefei, Zhu, Junyi, Wang, Wenyu, Yekhanin, Sergey
Generative modeling, representation learning, and classification are three core problems in machine learning (ML), yet their state-of-the-art (SoTA) solutions remain largely disjoint. In this paper, we ask: Can a unified principle address all three? Such unification could simplify ML pipelines and foster greater synergy across tasks. We introduce Latent Zoning Network (LZN) as a step toward this goal. At its core, LZN creates a shared Gaussian latent space that encodes information across all tasks. Each data type (e.g., images, text, labels) is equipped with an encoder that maps samples to disjoint latent zones, and a decoder that maps latents back to data. ML tasks are expressed as compositions of these encoders and decoders: for example, label-conditional image generation uses a label encoder and image decoder; image embedding uses an image encoder; classification uses an image encoder and label decoder. We demonstrate the promise of LZN in three increasingly complex scenarios: (1) LZN can enhance existing models (image generation): When combined with the SoTA Rectified Flow model, LZN improves FID on CIFAR10 from 2.76 to 2.59-without modifying the training objective. (2) LZN can solve tasks independently (representation learning): LZN can implement unsupervised representation learning without auxiliary loss functions, outperforming the seminal MoCo and SimCLR methods by 9.3% and 0.2%, respectively, on downstream linear classification on ImageNet. (3) LZN can solve multiple tasks simultaneously (joint generation and classification): With image and label encoders/decoders, LZN performs both tasks jointly by design, improving FID and achieving SoTA classification accuracy on CIFAR10. The code and trained models are available at https://github.com/microsoft/latent-zoning-networks. The project website is at https://zinanlin.me/blogs/latent_zoning_networks.html.