Large Language Model
Understanding Practitioners Perspectives on Monitoring Machine Learning Systems
Naveed, Hira, Grundy, John, Arora, Chetan, Khalajzadeh, Hourieh, Haggag, Omar
--Given the inherent non-deterministic nature of machine learning (ML) systems, their behavior in production environments can lead to unforeseen and potentially dangerous outcomes. For a timely detection of unwanted behavior and to prevent organizations from financial and reputational damage, monitoring these systems is essential. This paper explores the strategies, challenges, and improvement opportunities for monitoring ML systems from the practitioners' perspective. We conducted a global survey of 91 ML practitioners to collect diverse insights into current monitoring practices for ML systems. We aim to complement existing research through our qualitative and quantitative analyses, focusing on prevalent runtime issues, industrial monitoring and mitigation practices, key challenges, and desired enhancements in future monitoring tools. Our findings reveal that practitioners frequently struggle with runtime issues related to declining model performance, exceeding latency, and security violations. While most prefer automated monitoring for its increased efficiency, many still rely on manual approaches due to the complexity or lack of appropriate automation solutions. Practitioners report that the initial setup and configuration of monitoring tools is often complicated and challenging, particularly when integrating with ML systems and setting alert thresholds. Moreover, practitioners find that monitoring adds extra workload, strains resources, and causes alert fatigue. The desired improvements from the practitioners' perspective are: automated generation and deployment of monitors, improved support for performance and fairness monitoring, and recommendations for resolving runtime issues. These insights offer valuable guidance for the future development of ML monitoring tools that are better aligned with practitioners' needs. Machine Learning (ML) systems are being increasingly employed across various domains, including social media, e-commerce, and engineering - even critical domains such as finance, healthcare, and autonomous vehicles nowadays leverage ML to automate and enhance their services. Generative AI and Large Language Models (LLMs) have further boosted ML adoption by creating several new use cases [1], [2]. A typical ML system lifecycle begins by gathering requirements and preparing data, which is followed by the development of the ML component (experimentation, model training, and evaluation) and other traditional software components [3]. After development, the next step is integration and system testing. Once quality assurance is completed, the ML system is deployed to a production environment.
Devstral: Fine-tuning Language Models for Coding Agent Applications
Rastogi, Abhinav, Yang, Adam, Jiang, Albert Q., Liu, Alexander H., Sablayrolles, Alexandre, Hรฉliou, Amรฉlie, Martin, Amรฉlie, Agarwal, Anmol, Ehrenberg, Andy, Lo, Andy, Roux, Antoine, Darcet, Arthur, Mensch, Arthur, Bout, Baptiste, Roziรจre, Baptiste, De Monicault, Baudouin, Bamford, Chris, Wallenwein, Christian, Renaudin, Christophe, Lanfranchi, Clรฉmence, Denoix, Clรฉment, Barreau, Corentin, Mizelle, Darius Dabert Devon, Casas, Diego de las, Chane-Sane, Elliot, Fugier, Emilien, Hanna, Emma Bou, Berrada, Gabrielle, Delerce, Gauthier, Guinet, Gauthier, Novikov, Georgii, Neubig, Graham, Lample, Guillaume, Martin, Guillaume, Jaju, Himanshu, Ludziejewski, Jan, Rute, Jason, Delignon, Jean-Malo, Chabran, JeanHadrien, Studnia, Joachim, Barmentlo, Joep, Amar, Jonas, Roberts, Josselin Somerville, Denize, Julien, Saxena, Karan, Yadav, Karmesh, Khandelwal, Kartik, Chandu, Khyathi Raghavi, Jain, Kush, Lavaud, Lรฉlio Renard, Blier, Lรฉonard, Zhao, Lingxiao, Martin, Louis, Saulnier, Lucile, Gao, Luyu, Pellat, Marie, Guillaumin, Mathilde, Felardos, Mathis, Dinot, Matthieu, Darrin, Maxime, Augustin, Maximilian, Seznec, Mickaรซl, Gupta, Neha, Raghuraman, Nikhil, Duchenne, Olivier, Wang, Patricia, von Platen, Patrick, Saffer, Patryk, Jacob, Paul, Wambergue, Paul, Kurylowicz, Paula, Chagniot, Philomรจne, Stock, Pierre, Agrawal, Pravesh, Delacourt, Rรฉmi, Soletskyi, Roman, Sauvestre, Romain, Vaze, Sagar, Gandhi, Sanchit, Subramanian, Sandeep, Dalal, Shashwat, Gandhi, Siddharth, Ghosh, Soham, Mishra, Srijan, Aithal, Sumukh, Antoniak, Szymon, Scao, Teven Le, Lavril, Thibaut, Schueller, Thibault, Foubert, Thomas, Robert, Thomas, Wang, Thomas, Lacroix, Timothรฉe, Bewley, Tom, Nemychnikova, Valeriia, Paltz, Victor, Richard, Virgile, Li, Wen-Ding, Marshall, William, Wang, Xingyao, Zhang, Xuanyu, Wan, Yihan, Tang, Yunhao
We introduce Devstral-Small, a lightweight open source model for code agents with the best performance among models below 100B size. In this technical report, we give an overview of how we design and develop a model and craft specializations in agentic software development. The resulting model, Devstral-Small is a small 24B model, fast and easy to serve. Despite its size, Devstral-Small still attains competitive performance compared to models more than an order of magnitude larger.
TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language Models
Guan, Tong, Meng, Zijie, Li, Dianqi, Wang, Shiyu, Yang, Chao-Han Huck, Wen, Qingsong, Liu, Zuozhu, Siniscalchi, Sabato Marco, Jin, Ming, Pan, Shirui
Recent advances in multimodal time series learning underscore a paradigm shift from analytics centered on basic patterns toward advanced time series understanding and reasoning. However, existing multimodal time series datasets mostly remain at the level of surface alignment and question answering, without reaching the depth of genuine reasoning. The absence of well-defined tasks that genuinely require time series reasoning, along with the scarcity of high-quality data, has limited progress in building practical time series reasoning models (TSRMs). To this end, we introduce Time Series Reasoning Suite (TSR-Suite), which formalizes four atomic tasks that span three fundamental capabilities for reasoning with time series: (1) perception, acquired through scenario understanding and causality discovery; (2) extrapolation, realized via event-aware forecasting; and (3) decision-making, developed through deliberation over perception and extrapolation. TSR-Suite is the first comprehensive time series reasoning suite that supports not only thorough evaluation but also the data pipeline and training of TSRMs. It contains more than 23K samples, of which 2.3K are carefully curated through a human-guided hierarchical annotation process. Building on this foundation, we introduce TimeOmni-1, the first unified reasoning model designed to address diverse real-world problems demanding time series reasoning. The model is trained in multiple stages, integrating a mixture of task scenarios, novel reward functions, and tailored optimizations. Experiments show that TimeOmni-1 delivers strong out-of-distribution generalization across all tasks and achieves a high rate of valid responses. It significantly improves causality discovery accuracy (64.0% vs. 35.9% with GPT-4.1) and raises the valid response rate by over 6% compared to GPT-4.1 on the event-aware forecasting task.
OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment
Lin, Liang, Xu, Zhihao, Dong, Junhao, Zhao, Jian, Yuan, Yuchen, Zhang, Guibin, Yu, Miao, Zhang, Yiming, Yao, Zhengtao, Yi, Huahui, Liu, Dongrui, Li, Xinfeng, Wang, Kun
Large language model (LLM) alignment faces a critical dilemma when addressing multiple human preferences: improvements in one dimension frequently come at the expense of others, creating unavoidable trade-offs between competing objectives like helpfulness and harmlessness. While prior work mainly focuses on constraint-based optimization algorithms and data selection strategies to mitigate conflicts, these approaches overlook the fundamental issue of resolving conflicts directly at the parameter level. In this paper, we present OrthAlign, an innovative approach that pioneers a new paradigm by leveraging orthogonal subspace decomposition to fundamentally resolve gradient-level conflicts in multi-objective preference alignment. OrthAlign strategically decomposes parameter update spaces into orthogonal subspaces, ensuring that optimization toward different preferences occurs in mathematically non-interfering directions. Building upon this, we provide theoretical guarantees demonstrating that when parameter increments satisfy both orthogonal subspace constraints and spectral norm bounds, the resulting updates exhibit linear Lipschitz growth rather than exponential instability, ensuring stable convergence across all preference dimensions. Extensive experiments show that: I. OrthAlign achieves maximum single-preference improvements ranging from 34.61% to 50.89% after multiple-objective alignment across helpful, harmless, and truthful dimensions. II. With an average overall reward improvement of 13.96%.
Experience-Guided Reflective Co-Evolution of Prompts and Heuristics for Automatic Algorithm Design
Liu, Yihong, Li, Junyi, Zhao, Wayne Xin, Lu, Hongyu, Wen, Ji-Rong
Combinatorial optimization problems are traditionally tackled with handcrafted heuristic algorithms, which demand extensive domain expertise and significant implementation effort. Recent progress has highlighted the potential of automatic heuristics design powered by large language models (LLMs), enabling the automatic generation and refinement of heuristics. These approaches typically maintain a population of heuristics and employ LLMs as mutation operators to evolve them across generations. While effective, such methods often risk stagnating in local optima. To address this issue, we propose the Experience-Guided Reflective Co-Evolution of Prompt and Heuristics (EvoPH) for automatic algorithm design, a novel framework that integrates the island migration model with the elites selection algorithm to simulate diverse heuristics populations. In EvoPH, prompts are co-evolved with heuristic algorithms, guided by performance feedback. We evaluate our framework on two problems, i.e., Traveling Salesman Problem and Bin Packing Problem. Experimental results demonstrate that EvoPH achieves the lowest relative error against optimal solutions across both datasets, advancing the field of automatic algorithm design with LLMs.
UI-UG: A Unified MLLM for UI Understanding and Generation
Yang, Hao, Qiu, Weijie, Zhang, Ru, Fang, Zhou, Mao, Ruichao, Lin, Xiaoyu, Huang, Maji, Huang, Zhaosong, Guo, Teng, Liu, Shuoyang, Rao, Hai
Although Multimodal Large Language Models (MLLMs) have been widely applied across domains, they are still facing challenges in domain-specific tasks, such as User Interface (UI) understanding accuracy and UI generation quality. In this paper, we introduce UI-UG (a unified MLLM for UI Understanding and Generation), integrating both capabilities. For understanding tasks, we employ Supervised Fine-tuning (SFT) combined with Group Relative Policy Optimization (GRPO) to enhance fine-grained understanding on the modern complex UI data. For generation tasks, we further use Direct Preference Optimization (DPO) to make our model generate human-preferred UIs. In addition, we propose an industrially effective workflow, including the design of an LLM-friendly domain-specific language (DSL), training strategies, rendering processes, and evaluation metrics. In experiments, our model achieves state-of-the-art (SOT A) performance on understanding tasks, outperforming both larger general-purpose MLLMs and similarly-sized UI-specialized models. Our model is also on par with these larger MLLMs in UI generation performance at a fraction of the computational cost. We also demonstrate that integrating understanding and generation tasks can improve accuracy and quality for both tasks.
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
Wang, Junying, Zhang, Zicheng, Shen, Ye, Wu, Yalun, Liang, Yingji, Guo, Yijin, Wen, Farong, Li, Wenzhe, Zhao, Xuezhi, Jia, Qi, Zhai, Guangtao
High-quality, multi-modal benchmarks are crucial for advancing scientific reasoning in large models yet their manual creation is costly and unscalable. To address this bottleneck, we explore the potential for transforming Text-Only QA Pairs (TQAs) into high-quality Multi-Modal QA Pairs (MMQAs), which include three parts: 1) Task Definition \& Evaluation Rubric: We develop a TQA-to-MMQA framework and establish a comprehensive, multi-dimensional MMQA quality rubric that provides principles for the transformation. 2) Benchmark Construction: Then we construct two extensive benchmarks to rigorously evaluate state-of-the-art generation \& understanding models on the distinct tasks of MMQA generation \& MMQA quality evaluation. 3) Preliminary Solution: We develop an agentic system (Q-Mirror), which operationalizes our framework by integrating MMQA generation and evaluation into a closed loop for iterative refinement. Our experiments show that while state-of-the-art models can generate MMQAs, their outputs still leave substantial gaps, underscoring the need for reliable evaluation. We further demonstrate that top-tier understanding models align closely with human judgment in MMQA quality assessment. Leveraging both insights, the Q-Mirror agent raises average scores from 78.90 to 85.22 and pass rates from 72\% to 95\%, offering a practical path to large-scale scientific benchmarks.
Conda: Column-Normalized Adam for Training Large Language Models Faster
Wang, Junjie, Zhou, Pan, Dong, Yiming, Li, Huan, Li, Jia, Zhou, Xun, Lao, Qicheng, Fang, Cong, Lin, Zhouchen
Large language models (LLMs) have demonstrated impressive generalization and emergent capabilities, yet their pre-training remains computationally expensive and sensitive to optimization dynamics. While Adam-based optimizers offer fast convergence by adapting learning rates coordinate-wise, recent studies reveal that their updates often suffer from poor spectral conditioning and low-rank structures, hindering efficiency. Muon addresses this issue via global spectral normalization but lacks the per-coordinate adaptivity of Adam. In this work, we propose Column-Normalized Adam (Conda), a novel optimizer that bridges the strengths of both approaches. Conda projects updates into an orthogonal subspace and applies column-wise second moment normalization based on the projected gradients, thereby achieving both improved spectral conditioning and maintaining coordinate-wise adaptivity. This design alleviates the spectral pathologies of Adam while preserving its fast convergence behavior. Extensive experiments on the LLaMA and GPT-2 series show that Conda consistently outperforms AdamW, Muon, and other baselines in pre-training. Remarkably, on the LLaMA series, Conda achieves 2-2.5 the convergence speed of AdamW, measured in both training steps and training time. Further ablations demonstrate its robustness under diverse training setups. These results collectively highlight Conda as an effective and broadly applicable optimizer for large-scale LLM training. The code is released on https://github.com/jie040109/Conda
TENET: Leveraging Tests Beyond Validation for Code Generation
Hu, Yiran, Jiang, Nan, Liang, Shanchao, Wu, Yi, Tan, Lin
Test-Driven Development (TDD) is a widely adopted software engineering practice that requires developers to create and execute tests alongside code implementation, ensuring that software behavior is continuously validated and refined. In the era of vibe coding, where developers increasingly delegate code writing to large language models (LLMs) by specifying high-level intentions, TDD becomes even more crucial, as test cases serve as executable specifications that explicitly define and verify intended functionality beyond what natural-language descriptions and code context can convey. While vibe coding under TDD is promising, there are three main challenges: (1) selecting a small yet effective test suite to improve the generation accuracy and control the execution workload, (2) retrieving context such as relevant code effectively, and (3) systematically using test feedback for effective code refinement. To address these challenges, we introduce TENET, an LLM agent for generating functions in complex real-world repositories under the TDD setting. TENET features three components: (1) a novel test harness mechanism that selects a concise test suite to maximize diversity of target usage scenarios; (2) a tailored agent toolset that performs efficient retrieval of relevant code with interactive debugging; and (3) a reflection-based refinement workflow that iteratively analyzes failures, replenishes context, and applies code refinement. TENET achieves 69.08% and 81.77% Pass@1 on RepoCod and RepoEval benchmarks, outperforming the best agentic baselines by 9.49 and 2.17 percentage points, respectively. In addition, this is the first study of test-driven code generation with repository-level context, examining how different aspects of test suites affect the performance of LLM agents under the TDD setting.
Dual-Scale World Models for LLM Agents Towards Hard-Exploration Problems
LLM-based agents have seen promising advances, yet they are still limited in "hard-exploration" tasks requiring learning new knowledge through exploration. We present GLoW, a novel approach leveraging dual-scale world models, maintaining a trajectory frontier of high-value discoveries at the global scale, while learning from local trial-and-error in exploration through a Multi-path Advantage Reflection mechanism which infers advantage-based progress signals to guide exploration. To evaluate our framework for hard-exploration, we tackle the Jericho benchmark suite of text-based games, where GLoW achieves a new state-of-theart performance for LLM-based approaches. Compared to state-of-the-art RLbased methods, our approach achieves comparable performance while requiring 100-800x fewer environment interactions.