Country
Intel and AMD accused of allowing chips in Russian missiles
A woman and her relatives look at her home, which was damaged during a night of Russian missile and drone strikes, amid Russia's attack on Ukraine, in Novi Petrivtsi, outside Kyiv, on Saturday. Microchip manufacturers Intel, Advanced Micro Devices (AMD) and Texas Instruments were accused in a series of lawsuits of failing to keep their technology out of Russian-made weapons used to kill and wound civilians in Ukraine. Those companies -- along with a company owned by Warren Buffett's Berkshire Hathaway -- demonstrated willful ignorance" as third parties resold restricted chips to Russia to power drones and missiles in violation of U.S. sanctions, according to one of the five suits, filed Wednesday in state court in Texas. The lawsuits, filed on behalf of dozens of Ukrainian civilians by Mikal Watts and prominent law firm Baker & Hostetler, cite five attacks between 2023 and 2025 that killed dozens of people. One attack allegedly involved Iranian-made drones with components associated with Intel and AMD, while the others involved Russian-made KH-101 cruise missiles and Iskander ballistic missiles.
AI has entered the classroom - but is it the solution for overworked teachers?
AI has entered the classroom - but is it the solution for overworked teachers? Schools across the UK are trialling the use of deepfake teachers and even employing remote staff to deliver lessons hundreds of miles away from the classroom. It comes as the use of AI is becoming increasingly prevalent in schools. The government says AI has the power to transform education, and improve teacher workload, particularly around admin for teachers. The BBC has spoken to teachers, school leaders and unions who seem divided on what the future of the UK's classrooms should look like.
Revealed: Amazon Alexa's most-asked questions of 2025 - including 'how tall is Tom Cruise?' and 'how long do I poach an egg for?'
Ghislaine Maxwell's ultimate humiliation: Epstein's sex trafficker girlfriend poses in outrageous outfits and exposes herself in dozens of photos released from the billionaire paedophile's files I was falsely accused of being the Brown University shooter... Silent Trump flees growing storm over Epstein'cover-up' as he jets off for holidays without ANY comment Truth about THIS photo of Karoline Leavitt's face... and why if she was non-binary and disabled, Vanity Fair would never have done this: KENNEDY Why Conan O'Brien'stopped party guests calling 911' on Nick Reiner: Insiders reveal disturbing new details of final hours before Rob and Michele murders After 27 years as a TV anchor I was suddenly pulled off screens. My boss's explanation was a brutal lesson in loyalty Emily in Paris cast left'aghast' and'walking on eggshells' as off-camera drama becomes overwhelming... and whispers swirl about a CURSE Doctors said my hip pain was just tendinitis from sitting all day at work.
A Minimalist Optimizer Design for LLM Pretraining
Glentis, Athanasios, Li, Jiaxiang, Han, Andi, Hong, Mingyi
Training large language models (LLMs) typically relies on adaptive optimizers such as Adam, which introduce extra operations and require significant more memory to maintain first- and second-order moments than SGD. While recent works such as GaLore, Fira and APOLLO have proposed state-compressed variants to reduce memory consumption, a fundamental question remains: What are the minimum modifications to plain SGD needed to match state-of-the-art pretraining performance? We systematically investigate this question using a bottom-up approach, and identify two simple yet highly (memory- and compute-) efficient techniques: (1) column-wise gradient normalization (normalizing the gradient along the output dimension), which boosts SGD performance without momentum; and (2) applying first-order momentum only to the output layer, where gradient variance is highest. Combining these two techniques lead to SCALE (Stochastic Column-normAlized Last-layer momEntum), a simple optimizer for memory efficient pretraining. Across multiple LLaMA models (60M-1B), SCALE matches or exceeds the performance of Adam while using only 35-45% of the total memory. It also consistently outperforms memory-efficient optimizers such as GaLore, Fira and APOLLO, making it a strong candidate for large-scale pretraining under memory constraints. For LLaMA 7B model, SCALE outperforms the state-of-the-art memory-efficient methods APOLLO and Muon, in terms of both perplexity and memory consumption.
DeepMech: A Machine Learning Framework for Chemical Reaction Mechanism Prediction
Das, Manajit, Hoque, Ajnabiul, Baranwal, Mayank, Sunoj, Raghavan B.
Prediction of complete step-by-step chemical reaction mechanisms (CRMs) remains a major challenge. Whereas the traditional approaches in CRM tasks rely on expert-driven experiments or costly quantum chemical computations, contemporary deep learning (DL) alternatives ignore key intermediates and mechanistic steps and often suffer from hallucinations. We present DeepMech, an interpretable graph-based DL framework employing atom- and bond-level attention, guided by generalized templates of mechanistic operations (TMOps), to generate CRMs. Trained on our curated ReactMech dataset (~30K CRMs with 100K atom-mapped and mass-balanced elementary steps), DeepMech achieves 98.98+/-0.12% accuracy in predicting elementary steps and 95.94+/-0.21% in complete CRM tasks, besides maintaining high fidelity even in out-of-distribution scenarios as well as in predicting side and/or byproducts. Extension to multistep CRMs relevant to prebiotic chemistry, demonstrates the ability of DeepMech in effectively reconstructing 2 pathways from simple primordial substrates to complex biomolecules such as serine and aldopentose. Attention analysis identifies reactive atoms/bonds in line with chemical intuition, rendering our model interpretable and suitable for reaction design.
Supervised learning pays attention
Craig, Erin, Tibshirani, Robert
In-context learning with attention enables large neural networks to make context-specific predictions by selectively focusing on relevant examples. Here, we adapt this idea to supervised learning procedures such as lasso regression and gradient boosting, for tabular data. Our goals are to (1) flexibly fit personalized models for each prediction point and (2) retain model simplicity and interpretability. Our method fits a local model for each test observation by weighting the training data according to attention, a supervised similarity measure that emphasizes features and interactions that are predictive of the outcome. Attention weighting allows the method to adapt to heterogeneous data in a data-driven way, without requiring cluster or similarity pre-specification. Further, our approach is uniquely interpretable: for each test observation, we identify which features are most predictive and which training observations are most relevant. We then show how to use attention weighting for time series and spatial data, and we present a method for adapting pretrained tree-based models to distributional shift using attention-weighted residual corrections. Across real and simulated datasets, attention weighting improves predictive performance while preserving interpretability, and theory shows that attention-weighting linear models attain lower mean squared error than the standard linear model under mixture-of-models data-generating processes with known subgroup structure.
Provably Learning from Modern Language Models via Low Logit Rank
Golowich, Noah, Liu, Allen, Shetty, Abhishek
While modern language models and their inner workings are incredibly complex, recent work (Golowich, Liu & Shetty; 2025) has proposed a simple and potentially tractable abstraction for them through the observation that empirically, these language models all seem to have approximately low logit rank. Roughly, this means that a matrix formed by the model's log probabilities of various tokens conditioned on certain sequences of tokens is well approximated by a low rank matrix. In this paper, our focus is on understanding how this structure can be exploited algorithmically for obtaining provable learning guarantees. Since low logit rank models can encode hard-to-learn distributions such as noisy parities, we study a query learning model with logit queries that reflects the access model for common APIs. Our main result is an efficient algorithm for learning any approximately low logit rank model from queries. We emphasize that our structural assumption closely reflects the behavior that is empirically observed in modern language models. Thus, our result gives what we believe is the first end-to-end learning guarantee for a generative model that plausibly captures modern language models.
Drawback of Enforcing Equivariance and its Compensation via the Lens of Expressive Power
Chen, Yuzhu, Qin, Tian, Tian, Xinmei, He, Fengxiang, Tao, Dacheng
Equivariant neural networks encode symmetry as an inductive bias and have achieved strong empirical performance in wide domains. However, their expressive power remains not well understood. Focusing on 2-layer ReLU networks, this paper investigates the impact of equiv-ariance constraints on the expressivity of equivariant and layer-wise equivariant networks. By examining the boundary hyperplanes and the channel vectors of ReLU networks, we construct an example showing that equivariance constraints could strictly limit expressive power. However, we demonstrate that this drawback can be compensated via enlarging the model size. Furthermore, we show that despite a larger model size, the resulting architecture could still correspond to a hypothesis space with lower complexity, implying superior generalizability for equivariant networks.
Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search
Fadeeva, Ekaterina, Goloburda, Maiya, Rubashevskii, Aleksandr, Vashurin, Roman, Shelmanov, Artem, Nakov, Preslav, Sachan, Mrinmaya, Panov, Maxim
Consistency-based methods have emerged as an effective approach to uncertainty quantification (UQ) in large language models. These methods typically rely on several generations obtained via multinomial sampling, measuring their agreement level. However, in short-form QA, multinomial sampling is prone to producing duplicates due to peaked distributions, and its stochasticity introduces considerable variance in uncertainty estimates across runs. We introduce a new family of methods that employ beam search to generate candidates for consistency-based UQ, yielding improved performance and reduced variance compared to multinomial sampling. We also provide a theoretical lower bound on the beam set probability mass under which beam search achieves a smaller error than multinomial sampling. We empirically evaluate our approach on six QA datasets and find that its consistent improvements over multinomial sampling lead to state-of-the-art UQ performance.
Transformers for Tabular Data: A Training Perspective of Self-Attention via Optimal Transport
Candelieri, Antonio, Quadrio, Alessandro
This thesis examines self-attention training through the lens of Optimal Transport (OT) and develops an OT-based alternative for tabular classification. The study tracks intermediate projections of the self-attention layer during training and evaluates their evolution using discrete OT metrics, including Wasserstein distance, Monge gap, optimality, and efficiency. Experiments are conducted on classification tasks with two and three classes, as well as on a biomedical dataset. Results indicate that the final self-attention mapping often approximates the OT optimal coupling, yet the training trajectory remains inefficient. Pretraining the MLP section on synthetic data partially improves convergence but is sensitive to their initialization. To address these limitations, an OT-based algorithm is introduced: it generates class-specific dummy Gaussian distributions, computes an OT alignment with the data, and trains an MLP to generalize this mapping. The method achieves accuracy comparable to Transformers while reducing computational cost and scaling more efficiently under standardized inputs, though its performance depends on careful dummy-geometry design. All experiments and implementations are conducted in R.