Large Language Model
Learning a Zeroth-Order Optimizer for Fine-Tuning LLMs
Zhang, Kairun, Li, Haoyu, Zhao, Yanjun, Sun, Yifan, Zhang, Huan
Zeroth-order optimizers have recently emerged as a practical approach for fine-tuning large language models (LLMs), significantly reducing GPU memory consumption compared to traditional first-order methods. Yet, existing zeroth-order methods rely on hand-crafted, static sampling strategies that are not adaptable to model-specific structures. To address this, we propose ZO Fine-tuner, a learning-based zeroth-order optimizer for LLMs that automatically learns efficient perturbation strategies through a compact and memory-efficient design. Crucially, our approach is motivated by the observation that only a small number of foundation models and their derivatives are widely adopted in practice. Therefore, learning the optimizer once for a given LLM and reusing it across diverse downstream tasks is both feasible and highly desirable. Accordingly, ZO Fine-tuner is designed to scale learning to learn (L2L) to the foundation-model era by supporting one-time training per LLM with minimal overhead. Experiments on 4 LLMs and 7 datasets show that ZO Fine-tuner outperforms prior zeroth-order baselines in 82.1\% of task-model combinations, thereby demonstrating strong performance and scalability for efficient LLM fine-tuning. Our code is available at https://github.com/ASTRAL-Group/ZO_Fine_tuner.git.
VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
Li, Hengtao, Ding, Pengxiang, Suo, Runze, Wang, Yihao, Ge, Zirui, Zang, Dongyuan, Yu, Kexian, Sun, Mingyang, Zhang, Hongyin, Wang, Donglin, Su, Weihua
Vision-Language-Action (VLA) models enable embodied decision-making but rely heavily on imitation learning, leading to compounding errors and poor robustness under distribution shift. Reinforcement learning (RL) can mitigate these issues yet typically demands costly real-world interactions or suffers from sim-to-real gaps. We introduce VLA-RFT, a reinforcement fine-tuning framework that leverages a data-driven world model as a controllable simulator. Trained from real interaction data, the simulator predicts future visual observations conditioned on actions, allowing policy rollouts with dense, trajectory-level rewards derived from goal-achieving references. This design delivers an efficient and action-aligned learning signal, drastically lowering sample requirements. With fewer than 400 fine-tuning steps, VLA-RFT surpasses strong supervised baselines and achieves greater efficiency than simulator-based RL. Moreover, it exhibits strong robustness under perturbed conditions, sustaining stable task execution. Our results establish world-model-based RFT as a practical post-training paradigm to enhance the generalization and robustness of VLA models. For more details, please refer to https://vla-rft.github.io/.
Composer: A Search Framework for Hybrid Neural Architecture Design
Acun, Bilge, Sinha, Prasoon, Ardalani, Newsha, Bae, Sangmin, Golden, Alicia, Lin, Chien-Yu, Madhyastha, Meghana, Sun, Fei, Yadwadkar, Neeraja J., Wu, Carole-Jean
Hybrid model architectures that combine computational primitives (e.g., Attention, MLP) in different ratios have shown promising performance beyond Transformers. Some studies have shown that different interleavings of primitives can affect model quality as well. However, prior works explore the hybrid model architecture design space manually. Due to the large design space and training costs, discovering hybrid models that combine key computational primitives for pre-training is challenging. In this work, we take a principled approach in designing a modular hybrid model architecture search framework -- Composer. Composer explores model architectures at a small scale and extrapolates the top-performing model architectures to a larger scale using our proposed scaling strategies. Using Composer, we discover new hybrid LLM architectures that outperform Llama 3.2. Compared to Llama 3.2 and previous state-of-the-art baselines, the new model architectures consistently reduce validation loss at parameter scales of 350M-3B and improve evaluation accuracy on the downstream tasks by up to 2.8-8.3% (1.1-3.1% on average) while improving both training and inference efficiency.
Hollywood's Most Terrifying Nightmare Has Arrived
The Industry A.I. Is Ready to Crush Hollywood as We've Known It Generative video tools are ready to flood the market with robot actors and content--leaving studios and actors scrambling to catch up. Enter your email to receive alerts for this author. You can manage your newsletter subscriptions at any time. You're already subscribed to the aa_Nitish_Pahwa newsletter. You can manage your newsletter subscriptions at any time.
OpenAI's New Sora App Lets You Deepfake Yourself for Entertainment
OpenAI's latest app encourages users to generate a personal digital avatar and scroll AI-generated videos of themselves and their friends. On Tuesday, OpenAI released an AI video app called Sora . The platform is powered by OpenAI's latest video generation model, Sora 2, and revolves around a TikTok-like For You page of user-generated clips. This is the first product release from OpenAI that adds AI-generated sounds to videos. For now, it's available only on iOS and requires an invite code to join.
Exclusive: Mira Murati's Stealth AI Lab Launches Its First Product
Thinking Machines Lab, led by a group of prominent former OpenAI researchers, is betting that fine-tuning cutting-edge models will be the next frontier in AI. Thinking Machines Lab, a heavily funded startup cofounded by prominent researchers from OpenAI, has revealed its first product--a tool called Tinker that automates the creation of custom frontier AI models. "We believe [Tinker] will help empower researchers and developers to experiment with models and will make frontier capabilities much more accessible to all people," said Mira Murati, cofounder and CEO of Thinking Machines, in an interview with WIRED ahead of the announcement. Big companies and academic labs already fine-tune open source AI models to create new variants that are optimized for specific tasks, like solving math problems, drafting legal agreements, or answering medical questions. Typically, this work involves acquiring and managing clusters of GPUs and using various software tools to ensure that large-scale training runs are stable and efficient.
AI Is Learning to Do the Jobs of Doctors, Lawyers, and Consultants
RadVid-19, a program which identifies lung injuries through artificial intelligence, is used at the University of Sao Paulo in Brazil. RadVid-19, a program which identifies lung injuries through artificial intelligence, is used at the University of Sao Paulo in Brazil. The tasks resemble those that lawyers, doctors, financial analysts, and management consultants solve for a living. One asks for a diagnosis of a six-year-old patient based on nine pieces of multimedia evidence; another asks for legal advice on a musician's estate; a third calls for a valuation of part of a healthcare technology company. Mercor, which claims to supply "expert data" to every top AI company, says that it spent more than $500,000 to develop 200 tasks that test whether AIs can perform knowledge work with high economic value across law, medicine, finance, and management consulting.
Microsoft launches 365 Premium for consumers, retires Copilot Pro
When you purchase through links in our articles, we may earn a small commission. Microsoft 365 Premium combines the best of Microsoft Copilot with Microsoft 365. If you want your household to be on the cutting edge of AI, Microsoft has a deal for you: Microsoft 365 Premium, which combines the Family plan with Copilot Pro. At a new price of $19.99 per month, Microsoft 365 Premium sounds simple enough. Previously offered only to business users, it's now available to consumers as well.
Google's Gemini lets you chat with your smart home
When you purchase through links in our articles, we may earn a small commission. Google's Gemini lets you chat with your smart home With help from the revamped Google Home app, Gemini for Home promises to be the eyes and ears of your household. "Hey Google, set bedroom lamp to 50 percent." Such stilted voice commands have been the stuff of smart home for years, but with Gemini for Home, Google is promising a smart home you can have an actual conversation with. That idea--of a smart home that understands the big picture and can act with context in mind--underpins Google's ambitious Gemini for Home plans, which it's rolling out today following months of slow buildup.
The Download: OpenAI's caste bias problem, and how AI videos are made
The Download: OpenAI's caste bias problem, and how AI videos are made Plus: Taiwan has pushed back against America's chip request OpenAI is huge in India. Its models are steeped in caste bias. Caste bias is rampant in OpenAI's products, including ChatGPT, according to an MIT Technology Review investigation. Though CEO Sam Altman boasted about India being its second-largest market during the launch of GPT-5 in August, we found that both this new model, which now powers ChatGPT, as well as Sora, OpenAI's text-to-video generator, exhibit caste bias. This risks entrenching discriminatory views in ways that are currently going unaddressed. Mitigating caste bias in AI models is more pressing than ever.