Deep Learning
MMbeddings: Parameter-Efficient, Low-Overfitting Probabilistic Embeddings Inspired by Nonlinear Mixed Models
Simchoni, Giora, Rosset, Saharon
We present MMbeddings, a probabilistic embedding approach that reinterprets categorical embeddings through the lens of nonlinear mixed models, effectively bridging classical statistical theory with modern deep learning. By treating embeddings as latent random effects within a variational autoencoder framework, our method substantially decreases the number of parameters -- from the conventional embedding approach of cardinality $\times$ embedding dimension, which quickly becomes infeasible with large cardinalities, to a significantly smaller, cardinality-independent number determined primarily by the encoder architecture. This reduction dramatically mitigates overfitting and computational burden in high-cardinality settings. Extensive experiments on simulated and real datasets, encompassing collaborative filtering and tabular regression tasks using varied architectures, demonstrate that MMbeddings consistently outperforms traditional embeddings, underscoring its potential across diverse machine learning applications.
KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
Ye, Hancheng, Gao, Zhengqi, Ma, Mingyuan, Wang, Qinsi, Fu, Yuzhe, Chung, Ming-Yu, Lin, Yueqian, Liu, Zhijian, Zhang, Jianyi, Zhuo, Danyang, Chen, Yiran
Multi-agent large language model (LLM) systems are increasingly adopted for complex language processing tasks that require communication and coordination among agents. However, these systems often suffer substantial overhead from repeated reprocessing of overlapping contexts across agents. In typical pipelines, once an agent receives a message from its predecessor, the full context-including prior turns-must be reprocessed from scratch, leading to inefficient processing. While key-value (KV) caching is an effective solution for avoiding redundant computation in single-agent settings where prefixes remain unchanged, it cannot be directly reused in multi-agent scenarios due to diverging prefixes introduced by agent-specific context extensions. We identify that the core challenge lies in the offset variance of KV-caches across agents. To address this, we propose KVCOMM, a training-free framework that enables efficient prefilling in multi-agent inference by reusing KV-caches and aligning cache offsets of overlapping contexts under diverse prefix contexts. KVCOMM estimates and adjusts KV-caches for shared content by referencing a pool of cached examples-termed anchors-that store observed cache deviations under varying prefixes. The anchor pool is maintained and updated online, allowing dynamic adaptation to distinct user requests and context structures. KVCOMM achieves over 70% reuse rate across diverse multi-agent workloads, including retrieval-augmented generation, math reasoning, and collaborative coding tasks, all without quality degradation. Particularly, when each fully-connected agent receives 1K input tokens with 512 prefix tokens and 512 output tokens under a five-agent setting, KVCOMM achieves up to 7.8x speedup compared to the standard prefill pipeline, reducing TTFT from ~430 ms to ~55 ms.
OpenAI Signs 38 Billion Deal With Amazon
OpenAI has committed to buying billions of dollars worth of compute from AWS--the latest in a string of major deals brokered by the AI startup. OpenAI has signed a multi-year deal with Amazon to buy $38 billion worth of AWS cloud infrastructure to train its models and serve its users. The deal is yet another sign of the AI industry becoming increasingly entangled, with OpenAI now at the center of major partnerships with industry players including Google, Oracle, Nvidia, and AMD. The AWS agreement is also notable because OpenAI rose to prominence in part through its partnership with Microsoft--Amazon's biggest cloud rival. Amazon is also a major backer of one of OpenAI's key competitors, Anthropic.
OpenAI, Amazon sign 38bn AI deal
OpenAI has signed a new deal valued at $38bn with Amazon that will allow the artificial intelligence giant to run AI workloads across Amazon Web Services (AWS) cloud infrastructure. The seven-year deal announced on Monday is the first big AI push for the e-commerce giant after a restructuring last week. Experts say this does not mean that it will allow OpenAI to train its model on websites hosted by AWS - which includes the websites of The New York Times, Reddit and United Airlines. "Running OpenAI training inside AWS doesn't change their ability to scrape content from AWS-hosted websites [which they could already do for anything publicly readable]. This is strictly speaking about the economics of rent vs buy for GPU [graphics processing unit] capacity," Joshua McKenty, CEO of the AI detection company PolyguardAI, told Al Jazeera. The deal is also a major vote of confidence for the e-commerce giant's cloud unit, AWS, which some investors feared had fallen behind rivals Microsoft and Google in the artificial intelligence (AI) race.
OpenAI signs 38bn cloud computing deal with Amazon
OpenAI said the deal would give it access to hundreds of thousands of Nvidia graphics processors to train and run its AI models. OpenAI said the deal would give it access to hundreds of thousands of Nvidia graphics processors to train and run its AI models. Agreement to use AWS datacentres, and Nvidia chips inside them, part of $1.4tn spending spree on AI infrastructure Mon 3 Nov 2025 13.09 ESTLast modified on Mon 3 Nov 2025 15.16 EST OpenAI has signed a $38bn (รยฃ29bn) deal to use Amazon infrastructure to operate its artificial intelligence products, as part of a more than $1tn spending spree on computing power. The agreement with Amazon Web Services means OpenAI will be able to use AWS datacentres, and the Nvidia chips inside them, immediately. Last week, OpenAIรข s chief executive, Sam Altman, said his company had committed to spending $1.4tn on AI infrastructure, amid concerns over the sustainability of the boom in using and building datacentres.
ChatGPT owner OpenAI signs 38bn cloud computing deal with Amazon
OpenAI has signed a $38bn (ยฃ29bn) contract with Amazon to access its cloud computing infrastructure, as the start-up continues its run of major partnerships to secure computing power . In 2025, the ChatGPT maker has signed deals worth more than $1tn with Oracle, Broadcom, AMD and chip-making giant Nvidia. Its latest deal reduces its reliance on Microsoft. As part of the seven-year agreement, OpenAI will gain access to Nvidia graphics processors to train its artificial intelligence models. The deal follows a sweeping restructure of OpenAI last week which saw it convert away from being a non-profit and changed its relationship with Microsoft to give OpenAI more operational and financial freedom.
The Case That A.I. Is Thinking
The Case That A.I. Is Thinking ChatGPT does not have an inner life. Yet it seems to know what it's talking about. How convincing does the illusion of understanding have to be before you stop calling it an illusion? Dario Amodei, the C.E.O. of the artificial-intelligence company Anthropic, has been predicting that an A.I. "smarter than a Nobel Prize winner" in such fields as biology, math, engineering, and writing might come online by 2027. He envisions millions of copies of a model whirring away, each conducting its own research: a "country of geniuses in a datacenter." In June, Sam Altman, of OpenAI, wrote that the industry was on the cusp of building "digital superintelligence." "The 2030s are likely going to be wildly different from any time that has come before," he asserted. Meanwhile, the A.I. tools that most people currently interact with on a day-to-day basis are reminiscent of Clippy, the onetime Microsoft Office "assistant" that was actually more of a gadfly. A Zoom A.I. tool suggests that you ask it "What are some meeting icebreakers?" or instruct it to "Write a short message to share gratitude." Siri is good at setting reminders but not much else. A friend of mine saw a button in Gmail that said "Thank and tell anecdote." When he clicked it, Google's A.I. invented a funny story about a trip to Turkey that he never took. The rushed and uneven rollout of A.I. has created a fog in which it is tempting to conclude that there is nothing to see here--that it's all hype. There is, to be sure, plenty of hype: Amodei's timeline is science-fictional.
VeriFastScore: Speeding up long-form factuality evaluation
Rajendhran, Rishanth, Zadeh, Amir, Sarte, Matthew, Li, Chuan, Iyyer, Mohit
Metrics like FactScore and VeriScore that evaluate long-form factuality operate by decomposing an input response into atomic claims and then individually verifying each claim. While effective and interpretable, these methods incur numerous LLM calls and can take upwards of 100 seconds to evaluate a single response, limiting their practicality in large-scale evaluation and training scenarios. To address this, we propose VeriFastScore, which leverages synthetic data to fine-tune Llama3.1 8B for simultaneously extracting and verifying all verifiable claims within a given text based on evidence from Google Search. We show that this task cannot be solved via few-shot prompting with closed LLMs due to its complexity: the model receives ~4K tokens of evidence on average and needs to concurrently decompose claims, judge their verifiability, and verify them against noisy evidence. However, our fine-tuned VeriFastScore model demonstrates strong correlation with the original VeriScore pipeline at both the example level (r=0.80) and system level (r=0.94) while achieving an overall speedup of 6.6x (9.9x excluding evidence retrieval) over VeriScore. To facilitate future factuality research, we publicly release our VeriFastScore model and synthetic datasets.
Exploring Landscapes for Better Minima along Valleys
Zhao, Tong, Li, Jiacheng, Zhou, Yuanchang, Tan, Guangming, Jia, Weile
Finding lower and better-generalizing minima is crucial for deep learning. However, most existing optimizers stop searching the parameter space once they reach a local minimum. Given the complex geometric properties of the loss landscape, it is difficult to guarantee that such a point is the lowest or provides the best generalization. To address this, we propose an adaptor "E" for gradient-based optimizers. The adapted optimizer tends to continue exploring along landscape valleys (areas with low and nearly identical losses) in order to search for potentially better local minima even after reaching a local minimum. This approach increases the likelihood of finding a lower and flatter local minimum, which is often associated with better generalization. We also provide a proof of convergence for the adapted optimizers in both convex and non-convex scenarios for completeness. Finally, we demonstrate their effectiveness in an important but notoriously difficult training scenario, large-batch training, where Lamb is the benchmark optimizer. Our testing results show that the adapted Lamb, ALTO, increases the test accuracy (generalization) of the current state-of-the-art optimizer by an average of 2.5% across a variety of large-batch training tasks. This work potentially opens a new research direction in the design of optimization algorithms.