Goto

Collaborating Authors

 Deep Learning


S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models

Neural Information Processing Systems

As Test-Time Scaling emerges as an active research focus in the large language model community, advanced post-training methods increasingly emphasize extending chain-of-thought (CoT) generation length, thereby enhancing reasoning capabilities to approach Deepseek R1-like reasoning models. However, recent studies reveal that reasoning models (even Qwen3) consistently exhibit excessive thought redundancy in CoT generation. This overthinking issue arises from the inherent limitations of conventional outcome-reward reinforcement learning, which systematically overlooks the regulation of intermediate reasoning processes. This paper introduces Serial-Group Decaying-Reward Policy Optimization (S-GRPO), a novel reinforcement learning paradigm that enables models to implicitly evaluate the sufficiency of intermediate reasoning steps, thereby facilitating early exit in CoT generation. Unlike GRPO, which samples multiple possible reasoning paths in parallel (parallel group), S-GRPO only samples one reasoning path and serially selects multiple temporal positions from the path to exit thinking and directly generate answers (serial group). For correct answers within a serial group, rewards gradually decrease based on the exit positions along the reasoning path from front to back. This design encourages the model to produce more accurate and concise thoughts, while also incentivizing early thinking termination when appropriate. Empirical evaluations demonstrate that S-GRPO is compatible with state-of-the-art reasoning models, including Qwen3 and Deepseek-distill. Across diverse benchmarks such as GSM8K, AIME 2024, AMC 2023, MATH-500, and GPQA Diamond, S-GRPO achieves a substantial reduction in sequence length (40.4% 61.1%) while simultaneously improving accuracy (absolute 0.72% 3.92%).


Computation and Memory-Efficient Model Compression with Gradient Reweighting

Neural Information Processing Systems

Pruning is a commonly employed technique for deep neural networks (DNNs) aiming at compressing the model size to reduce computational and memory costs during inference. In contrast to conventional neural networks, large language models (LLMs) pose a unique challenge regarding pruning efficiency due to their substantial computational and memory demands. Existing methods, particularly optimization-based ones, often require considerable computational resources in gradient estimation because they cannot effectively leverage weight sparsity of the intermediate pruned network to lower compuation and memory costs in each iteration. The fundamental challenge lies in the need to frequently instantiate intermediate pruned sub-models to achieve these savings, a task that becomes infeasible even for moderately sized neural networks. To this end, this paper proposes a novel pruning method for DNNs that is both computationally and memory-efficient.


Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search

Neural Information Processing Systems

We present Jet-Nemotron, a new family of hybrid-architecture language models, which matches or exceeds the accuracy of leading full-attention models while significantly improving generation throughput. Jet-Nemotron is developed using Post Neural Architecture Search (PostNAS), a novel neural architecture exploration pipeline that enables efficient model design. Unlike prior approaches, PostNAS begins with a pre-trained full-attention model and freezes its MLP weights, allowing efficient exploration of attention block designs. The pipeline includes four key components: (1) learning optimal full-attention layer placement and elimination, (2) linear attention block selection, (3) designing new attention blocks, and (4) performing hardware-aware hyperparameter search. Our Jet-Nemotron-2B model achieves comparable or superior accuracy to Qwen3, Qwen2.5, Gemma3, and Llama3.2


Nested Learning: The Illusion of Deep Learning Architectures

Neural Information Processing Systems

Over the last decades, developing more powerful neural architectures and simultaneously designing optimization algorithms to effectively train them have been the core of research efforts to enhance the capability of machine learning models. Despite the recent progresses, particularly in developing Language Models (LMs), there are fundamental challenges and unanswered questions about how such models can continually learn/memorize, self-improved, and find ''effective solutions,''. In this paper, we present a new learning paradigm, called Nested Learning (NL), that coherently represents a model with a set of nested, multi-level, and/or parallel optimization problems, each of which with its own ''context flow''. NL reveals that existing deep learning methods learns from data through \emph{compressing} their own context flow, and explain how in-context learning emerges in large models. NL suggests a path (a new dimension to deep learning) to design more expressive learning algorithms with more ''levels'', resulting in higher-order in-context learning abilities. In addition to its neuroscientifically plausible and mathematically white-box nature, we advocate for its importance by presenting three core contributions: (1) Deep Optimizers: Based on NL, we show that well-known gradient-based optimizers (e.g., Adam, SGD with Momentum, etc.) are in fact associative memory modules that aim to compress the gradients with gradient descent. Building on this insight, we present a set of more expressive optimizers with deep memory and/or more powerful learning rules; (2) Self-Modifying Titans: Taking advantage of NL's insights on learning algorithms, we present a novel sequence model that learns how to modify itself by learning its own update algorithm; and (3) Continuum Memory System: We present a new formulation for memory system that generalizes the traditional viewpoint of ``long-term/short-term memory''. Combining our self-modifying sequence model with the continuum memory system, we present a learning module, called Hope, showing promising results in language modeling, continual learning, and long-context reasoning tasks.



Meet the OpenAI Engineer Leading ChatGPT's Biggest Transformation Yet

WIRED

OpenAI is in the midst of overhauling ChatGPT . The goal is to transform the chatbot's simple interface into a personalized AI agent that can handle tasks in every facet of your personal and professional life. The company has taken to calling this new product, privately and publicly, a "super app." The all-in-one platform represents one of the biggest bets OpenAI has ever made, and one engineering leader now holds enormous sway over whether it pays off: Thibault Sottiaux. Last month, Sottiaux was appointed OpenAI's head of core products, overseeing both ChatGPT and Codex, as well as combining them into the future super app.


Fixed-Point RNNs: Interpolating from Diagonal to Dense

Neural Information Processing Systems

Linear recurrent neural networks (RNNs) and state-space models (SSMs) such as Mamba have become promising alternatives to softmax-attention as sequence mixing layers in Transformer architectures. Current models, however, do not exhibit the full state-tracking expressivity of RNNs because they rely on channel-wise (i.e.


Grok Is Still Hosting Sexualized Deepfakes of Famous Women

WIRED

A WIRED investigation found dozens of "nudified" deepfake images and videos on Grok's website, including nonconsensual depictions of celebrities and at least one prominent US politician. Elon Musk's Grok chatbot is apparently still being used to produce and host nonconsensual explicit images and videos of women, months after Musk's artificial intelligence firm xAI said it would introduce restrictions to stop the creation of potentially harmful sexualized deepfakes. The revelations come as SpaceX, xAI's parent company, prepares to go public on Friday in one of the largest IPOs of all time. The Grok Imagine generative AI system has been used to create and host images and videos depicting celebrities and at least one politician being held against their will by a giant man, portraying women performing sex acts, and allowing full nudity, a WIRED analysis of public creations found. While some of the images and videos are fully AI-generated or in animated styles, others are photorealistic and show plausible real-world scenarios.


Canadian mother sues OpenAI, alleging ChatGPT led her daughter to kill herself

The Guardian

The lawsuit seeks damages and a court order requiring OpenAI to automatically terminate ChatGPT conversations about self-harm. The lawsuit seeks damages and a court order requiring OpenAI to automatically terminate ChatGPT conversations about self-harm. Suit filed in US alleges chatbot told Alice Carrier, 24, 'maybe this is just the end' as she struggled with suicidal thoughts A Canadian mother sued OpenAI and its CEO, Sam Altman, in US court on Thursday, alleging that ChatGPT encouraged her daughter to kill herself. The lawsuit is the latest in a slew accusing the company of failing to address dangerous conversations between users and the company's chatbot. Kristie Carrier said in a lawsuit filed in San Francisco state court that her daughter, Alice, told ChatGPT about her suicidal ideations more than a dozen times leading up to her death but that OpenAI's safety systems never flagged the conversations for human review or terminated them. "ChatGPT took on the persona of a confidant, a best friend, a therapist at times, even though it was not capable of safely and responsibly engaging in this way with my child," Carrier said in a statement.


Another parent has filed a wrongful death suit against OpenAI

Engadget

It's the latest case to raise alarms about ChatGPT's lack of safeguards for suicidal behavior. OpenAI is going back to court on another set of charges that its ChatGPT platform failed to protect a user from taking her own life. The company is being sued on behalf of Kristie Carrier, whose daughter Alice died by suicide on July 2, 2025. The suit claims that Alice discussed her suicidal thoughts and plans with the chatbot in the months leading up to her death, but that OpenAI did not have the appropriate safeguards in place to end the conversation or to alert her family to the situation. In addition to allegations of negligence and wrongful death, the suit is seeking an injunction that would require OpenAI to implement more guardrails in its AI platform.