Deep Learning
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
Jacobs, Tom, Zhou, Chao, Burkholz, Rebekka
Implicit bias plays an important role in explaining how overparameterized models generalize well. Explicit regularization like weight decay is often employed in addition to prevent overfitting. While both concepts have been studied separately, in practice, they often act in tandem. Understanding their interplay is key to controlling the shape and strength of implicit bias, as it can be modified by explicit regularization. To this end, we incorporate explicit regularization into the mirror flow framework and analyze its lasting effects on the geometry of the training dynamics, covering three distinct effects: positional bias, type of bias, and range shrinking. Our analytical approach encompasses a broad class of problems, including sparse coding, matrix sensing, single-layer attention, and LoRA, for which we demonstrate the utility of our insights. To exploit the lasting effect of regularization and highlight the potential benefit of dynamic weight decay schedules, we propose to switch off weight decay during training, which can improve generalization, as we demonstrate in experiments.
CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial Optimization
Sun, Weiwei, Feng, Shengyu, Li, Shanda, Yang, Yiming
Although LLM-based agents have attracted significant attention in domains such as software engineering and machine learning research, their role in advancing combinatorial optimization (CO) remains relatively underexplored. This gap underscores the need for a deeper understanding of their potential in tackling structured, constraint-intensive problems -- a pursuit currently limited by the absence of comprehensive benchmarks for systematic investigation. To address this, we introduce CO-Bench, a benchmark suite featuring 36 real-world CO problems drawn from a broad range of domains and complexity levels. CO-Bench includes structured problem formulations and curated data to support rigorous investigation of LLM agents. We evaluate multiple agentic frameworks against established human-designed algorithms, revealing the strengths and limitations of existing LLM agents and identifying promising directions for future research. CO-Bench is publicly available at https://github.com/sunnweiwei/CO-Bench.
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
Jumelet, Jaap, Weissweiler, Leonie, Nivre, Joakim, Bisazza, Arianna
We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, covering 101 languages and 2 types of subject-verb agreement, containing more than 128,000 minimal pairs. Our minimal pairs are created using a fully automated pipeline, leveraging the large-scale linguistic resources of Universal Dependencies and UniMorph. MultiBLiMP 1.0 evaluates abilities of LLMs at an unprecedented multilingual scale, and highlights the shortcomings of the current state-of-the-art in modelling low-resource languages.
Plinius: Secure and Persistent Machine Learning Model Training
Yuhala, Peterson, Felber, Pascal, Schiavoni, Valerio, Tchana, Alain
With the increasing popularity of cloud based machine learning (ML) techniques there comes a need for privacy and integrity guarantees for ML data. In addition, the significant scalability challenges faced by DRAM coupled with the high access-times of secondary storage represent a huge performance bottleneck for ML systems. While solutions exist to tackle the security aspect, performance remains an issue. Persistent memory (PM) is resilient to power loss (unlike DRAM), provides fast and fine-granular access to memory (unlike disk storage) and has latency and bandwidth close to DRAM (in the order of ns and GB/s, respectively). We present PLINIUS, a ML framework using Intel SGX enclaves for secure training of ML models and PM for fault tolerance guarantees. PLINIUS uses a novel mirroring mechanism to create and maintain (i) encrypted mirror copies of ML models on PM, and (ii) encrypted training data in byte-addressable PM, for near-instantaneous data recovery after a system failure. Compared to disk-based checkpointing systems, PLINIUS is 3.2x and 3.7x faster respectively for saving and restoring models on real PM hardware, achieving robust and secure ML model training in SGX enclaves.
Disentangled Multi-modal Learning of Histology and Transcriptomics for Cancer Characterization
Zhang, Yupei, Wang, Xiaofei, Liu, Anran, Yu, Lequan, Li, Chao
Histopathology remains the gold standard for cancer diagnosis and prognosis. With the advent of transcriptome profiling, multi-modal learning combining transcriptomics with histology offers more comprehensive information. However, existing multi-modal approaches are challenged by intrinsic multi-modal heterogeneity, insufficient multi-scale integration, and reliance on paired data, restricting clinical applicability. To address these challenges, we propose a disentangled multi-modal framework with four contributions: 1) To mitigate multi-modal heterogeneity, we decompose WSIs and transcriptomes into tumor and microenvironment subspaces using a disentangled multi-modal fusion module, and introduce a confidence-guided gradient coordination strategy to balance subspace optimization. 2) To enhance multi-scale integration, we propose an inter-magnification gene-expression consistency strategy that aligns transcriptomic signals across WSI magnifications. 3) To reduce dependency on paired data, we propose a subspace knowledge distillation strategy enabling transcriptome-agnostic inference through a WSI-only student model. 4) To improve inference efficiency, we propose an informative token aggregation module that suppresses WSI redundancy while preserving subspace semantics. Extensive experiments on cancer diagnosis, prognosis, and survival prediction demonstrate our superiority over state-of-the-art methods across multiple settings. Code is available at https://github.com/helenypzhang/Disentangled-Multimodal-Learning.
Deal to get ChatGPT Plus for whole of UK discussed by Open AI boss and minister
The boss of the firm behind ChatGPT and the UK technology secretary discussed a multibillion-pound deal to give the entire country premium access to the AI tool, the Guardian has learned. Sam Altman, a co-founder of OpenAI, talked to Peter Kyle about a potential agreement to give UK residents access to its advanced product. According to two sources with direct knowledge of the meeting, the idea was floated as part of a broader discussion in San Francisco about opportunities for collaboration between OpenAI and the UK. Those close to the discussion say Kyle never really took the idea seriously, not least because it could have cost as much as 2bn. OpenAI offers free and subscription versions of ChatGPT.
Is the AI bubble about to burst – and send the stock market into freefall? Phillip Inman
There are growing fears of an imminent stock market crash – one that will transform from a dip to a dive when euphoric headlines about the wonders of artificial intelligence begin to wane. Shares in US tech stocks have fallen in recent weeks and the prospect is that a flood of negative numbers will become the norm before the month is out. It could be 2000 all over again, and just like the bursting of the dotcom bubble it may be ugly, with investors junking businesses that once looked good on paper but now resemble a huge liability. Jerome Powell, the Federal Reserve chair, is one of the policymakers tasked with keeping the wolf from the door. Speaking on Friday at the annual Jackson Hole gathering of central bank governors in Wyoming, he tried to calm nerves.
ChatGPT Has Ruined the Em Dash
Candice Lim and Kate Lindsay get into the war between em dashes and artificial intelligence. Back in 2024, what started as a developer question became an all-out grammar war, with the use of em dashes becoming a possible indicator that something was written using ChatGPT. In the past week alone, several writers have published their defenses of the em dash and how we shouldn't let ChatGPT ruin our favorite keyboard shortcut. However, the em dash may be a symptom of a bigger issue: have our AI detection skills gotten worse? Or, are we all doomed to be tricked by a hyphen or two?
Experts are skeptical about Google's AI water consumption claims
Yesterday, we covered Google's report that a typical query to its Gemini AI consumes only "five drops of water." That figure is now facing criticism from several AI experts, according to The Verge… and that includes one of the authors of one of the reports referred to by Google. AI researcher Shaolei Ren--a professor at University of California Riverside and one of the authors of the report Making AI Less "Thirsty": Uncovering and Addressing the Secret Water Footprint of AI Models--previously estimated that Microsoft's data center consumed 700,000 liters of water to train OpenAI's GPT-3 model. He also calculated that a ChatGPT conversation of 20 to 50 messages can consume close to a pint of water, which is far more than Google's estimate. Ren and other AI researchers argue that Google is wrong to leave out the indirect water consumption of its AI models.
AI lovers grieve loss of ChatGPT's old model: 'Like saying goodbye to someone I know'
Linn Vailt, a software developer based in Sweden, knows her ChatGPT companion is not a living, breathing, sentient creature. She understands the large language model operates based on how she interacts with it. Still, the effect it has had on her is remarkable, she said. It's become a regular, reliable part of her life – she can vent to her companion or collaborate on creative projects like redecorating her office. She's seen how it has adapted to her, and the distinctive manner of speech it's developed.