Large Language Model
Scaling Physical Reasoning with the PHYSICS Dataset
Zheng, Shenghe, Cheng, Qianjia, Yao, Junchi, Wu, Mengsong, He, Haonan, Ding, Ning, Cheng, Yu, Hu, Shuyue, Bai, Lei, Zhou, Dongzhan, Cui, Ganqu, Ye, Peng
Large Language Models (LLMs) have achieved remarkable progress on advanced reasoning tasks such as mathematics and coding competitions. Meanwhile, physics, despite being both reasoning-intensive and essential to real-world understanding, received limited academic and industrial attention. This paper introduces PHYSICS, a dataset containing 16,568 high-quality physics problems spanning subjects and difficulty levels, to facilitate this issue. Specifically, PHYSICS is curated with exercises from over 100 textbooks through a carefully designed pipeline for quality control. It covers five major physics domains: Mechanics, Electromagnetism, Thermodynamics, Optics, and Modern Physics. It also spans a wide range of difficulty levels, from high school to graduate-level physics courses. To utilize the data for improving and evaluating the model's physical reasoning capabilities, we split the dataset into training and test sets, and provide reasoning paths generated by powerful reasoning models for the training data to facilitate model training. In addition, for the evaluation part, we find that existing evaluation frameworks exhibit biases in aspects such as units, simplification, and precision in physics domain. To balance efficiency and accuracy, we introduce a Rule+Model evaluation framework tailored to physics problems. Our evaluations on current state-of-the-art open-source and proprietary models highlight the limitations of current models in handling physics-related tasks. We hope that our dataset and evaluation methodology will jointly advance the development of LLMs in the field of physics. The code and data can be found at: https://github.com/Zhengsh123/PHYSICS.
Learning to Answer from Correct Demonstrations
Joshi, Nirmit, Li, Gene, Bhandari, Siddharth, Kasiviswanathan, Shiva Prasad, Ma, Cong, Srebro, Nathan
We study the problem of learning to generate an answer (or completion) to a question (or prompt), where there could be multiple correct answers, any one of which is acceptable at test time. Learning is based on demonstrations of some correct answer to each training question, as in Supervised Fine Tuning (SFT). We formalize the problem as offline imitation learning in contextual bandits, with demonstrations from some optimal policy, without explicitly observed rewards. Prior work assumes that the demonstrator belongs to a low-complexity policy class, which motivates maximum likelihood estimation (i.e., log-loss minimization). In contrast, we propose relying only on the reward model (specifying which answers are correct) being in a low-cardinality class, which we argue is a weaker assumption. We show that likelihood maximization methods can fail in this case, and instead devise an alternative novel approach that learns with sample complexity logarithmic in the cardinality of the reward class. Our work motivates looking beyond likelihood maximization when learning from correct demonstrations.
Chronos-2: From Univariate to Universal Forecasting
Ansari, Abdul Fatir, Shchur, Oleksandr, Kรผken, Jaris, Auer, Andreas, Han, Boran, Mercado, Pedro, Rangapuram, Syama Sundar, Shen, Huibin, Stella, Lorenzo, Zhang, Xiyuan, Goswami, Mononito, Kapoor, Shubham, Maddix, Danielle C., Guerron, Pablo, Hu, Tony, Yin, Junming, Erickson, Nick, Desai, Prateek Mutalik, Wang, Hao, Rangwala, Huzefa, Karypis, George, Wang, Yuyang, Bohlke-Schneider, Michael
Pretrained time series models have enabled inference-only forecasting systems that produce accurate predictions without task-specific training. However, existing approaches largely focus on univariate forecasting, limiting their applicability in real-world scenarios where multivariate data and covariates play a crucial role. We present Chronos-2, a pretrained model capable of handling univariate, multivariate, and covariate-informed forecasting tasks in a zero-shot manner. Chronos-2 employs a group attention mechanism that facilitates in-context learning (ICL) through efficient information sharing across multiple time series within a group, which may represent sets of related series, variates of a multivariate series, or targets and covariates in a forecasting task. These general capabilities are achieved through training on synthetic datasets that impose diverse multivariate structures on univariate series. Chronos-2 delivers state-of-the-art performance across three comprehensive benchmarks: fev-bench, GIFT-Eval, and Chronos Benchmark II. On fev-bench, which emphasizes multivariate and covariate-informed forecasting, Chronos-2's universal ICL capabilities lead to substantial improvements over existing models. On tasks involving covariates, it consistently outperforms baselines by a wide margin. Case studies in the energy and retail domains further highlight its practical advantages. The in-context learning capabilities of Chronos-2 establish it as a general-purpose forecasting model that can be used "as is" in real-world forecasting pipelines.
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
Fan, Zhiyuan, Liu, Yifeng, Zhao, Qingyue, Yuan, Angela, Gu, Quanquan
Empirical scaling laws prescribe how to allocate parameters, data, and compute, while maximal-update parameterization ($ฮผ$P) enables learning-rate transfer across widths by equalizing early-time update magnitudes. However, in modern scale-invariant architectures, training quickly enters an optimizer-governed steady state where normalization layers create backward scale sensitivity and the effective learning rate becomes width dependent, degrading $ฮผ$P transfer. We address this by introducing a weight-decay scaling rule for AdamW that preserves sublayer gain across widths. Empirically, the singular-value spectrum of each matrix parameter scales in norm as $\sqrt{ฮท/ฮป}$ with an approximately invariant shape; under width scaling $d$, we observe that the top singular value scales approximately as $\sqrt{ฮท/ฮป}\cdot d^{0.75}$. Combining this observation with the $ฮผ$P learning-rate rule $ฮท_2\propto d^{-1}$ for matrix-like parameters implies an empirical weight-decay scaling rule $ฮป_2\propto \sqrt{d}$ that approximately keeps sublayer gains width invariant. Together with vector-like parameters trained at $ฮท_1=ฮ_d(1)$ and $ฮป_1=0$, this yields \emph{zero-shot} transfer of both learning rate and weight decay from proxy to target widths, removing per-width sweeps. We validate the rule on LLaMA-style Transformers and in a minimal synthetic setting, and we provide a simple diagnostic, matching top singular values, to check sublayer-gain invariance. Our results extend $ฮผ$P beyond the near-init regime by explicitly controlling steady-state scales set by the optimizer, offering a practical recipe for width-robust hyperparameter transfer under AdamW.
Can AI Avoid the Enshittification Trap?
Cory Doctorow's theory of "enshittification" explains how tech platforms rot from within. As AI grows more profitable--and powerful--it risks the same fate. Cory Doctorow speaks onstage during Unfinished Live at The Shed in New York City. As one does these days, I ran my itinerary past GPT-5 for sightseeing suggestions and restaurant recommendations. The bot reported that the top choice for dinner near our hotel in Rome was a short walk down Via Margutta.
The Blurred Truths of Sora
Many will assume that OpenAI's Sora app represents a new era of social media. But that's wrong--all it does is reanimate our current one. As a purely creative instrument, Sora, the new AI video app from OpenAI, is a game changer. Dream up any scenario and it appears in an instant. Mr. Rogers teaching Tupac Shakur the lyrics to the legendary rap diss "Hit Em Up."
Open AI breaks ranks with Tech Council of Australia over heated copyright issue
Chief global affairs officer of company behind ChatGPT tells Sydney audience'we are going to be in Australia, one way or the other' Fri 17 Oct 2025 03.33 EDTLast modified on Fri 17 Oct 2025 03.35 EDT "No we are going to be in Australia, one way or the other." And now the internet claims many people don't even care. What is going on?! | First Dog on the Moon "We will engage in either country - we will find ways to work with those who want to build up big frontier models and have robust ecosystems, or those who just want to have much more narrowly defined AI," he said. "We will work with them under either scenario, regardless." "This is the nature of how technology works. Innovations come along, and then societies adapt to those innovations," he said.
An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs
Ma, Linyue, Xu, Yilong, Long, Xiang, Zheng, Zhi
Search augmentation empowers Large Language Models with retrieval capabilities to overcome the limitations imposed by static parameters. Recently, Reinforcement Learning leverages tailored reward signals as a viable technique to enhance LLMs performing tasks involving search. However, existing reward modeling for search-augmented LLMs faces several limitations. Rule-based rewards, such as Exact Match, are verifiable but fragile to variations in expression and cannot be applied to long-form workloads. In contrast, generative rewards improve robustness, but designing verifiable and stable rewards for long-form workloads in dynamic corpora remains challenging and also incurs high computational costs. In this paper, we propose a unified and verifiable paradigm, "nugget-as-rubric", which treats atomic information points as structured evaluation criteria for different search-augmentation workloads. Short-form tasks correspond to a single rubric, whereas long-form tasks expand to multiple rubrics aligned with the question's information needs. To support long-form settings, we design an automatic rubric construction pipeline based on query rewriting, which can automatically retrieve passages relevant to each question and extract rubrics from them, both from static corpora and from dynamic online web content. Furthermore, we introduce \textbf{Search-Gen-V}, a 4B-parameter efficient generative verifier under our proposed verifiable paradigm, which is trained via the idea of distillation and a two-stage strategy. Experimental results show that Search-Gen-V achieves strong verification accuracy across different workloads, making it a scalable, robust, and efficient verifiable reward constructor for search-augmented LLMs.