Large Language Model
Apriel-Nemotron-15B-Thinker
Radhakrishna, Shruthan, Parikh, Soham, Sarda, Gopal, Turkkan, Anil, Vohra, Quaizar, Li, Raymond, Jhamb, Dhruv, Ogueji, Kelechi, Shukla, Aanjaneya, Bamgbose, Oluwanifemi, Liang, Toby, Kumar, Luke, Ostapenko, Oleksiy, Malay, Shiva Krishna Reddy, Tiwari, Aman, Bogavelli, Tara, Yadav, Vikas, Mehta, Jash, Mittal, Saloni, Kalkunte, Akshay, Pattnaik, Pulkit, Slimi, Khalil, Sreeram, Anirudh, Nair, Jishnu, Oladipo, Akintunde, Maiya, Shashank, Mahajan, Khyati, Maheshwary, Rishabh, Hashemi, Masoud, Mudumba, Sai Rajeswar, Madhusudhan, Sathwik Tejaswi, Scholak, Torsten, Paquet, Sebastien, Davasam, Sagar, Sunkara, Srinivas
While large language models (LLMs) have achieved remarkable reasoning capabilities across domains like code, math and other enterprise tasks, their significant memory and computational costs often preclude their use in practical enterprise settings. To this end, we introduce Apriel-Nemotron-15B-Thinker, a 15-billion parameter model in the ServiceNow Apriel SLM series that achieves performance against medium sized state-of-the-art models such as o1-mini, QWQ32B, and EXAONE-Deep-32B while maintaining only half the memory footprint of those alternatives. Apriel-Nemotron-15B-Thinker model is trained in a four stage training pipeline including 1) Base Model upscaling, 2) Continual Pre-training 3) Supervised Fine-tuning (SFT) and 4) Reinforcement Learning using GRPO. Comprehensive evaluations across a diverse suite of benchmarks consistently demonstrate that our Apriel-Nemotron-15B-Thinker model matches or exceeds the performance of its 32-billion parameter counterparts, despite being less than half their size.
gpt-oss-120b & gpt-oss-20b Model Card
OpenAI, null, :, null, Agarwal, Sandhini, Ahmad, Lama, Ai, Jason, Altman, Sam, Applebaum, Andy, Arbus, Edwin, Arora, Rahul K., Bai, Yu, Baker, Bowen, Bao, Haiming, Barak, Boaz, Bennett, Ally, Bertao, Tyler, Brett, Nivedita, Brevdo, Eugene, Brockman, Greg, Bubeck, Sebastien, Chang, Che, Chen, Kai, Chen, Mark, Cheung, Enoch, Clark, Aidan, Cook, Dan, Dukhan, Marat, Dvorak, Casey, Fives, Kevin, Fomenko, Vlad, Garipov, Timur, Georgiev, Kristian, Glaese, Mia, Gogineni, Tarun, Goucher, Adam, Gross, Lukas, Guzman, Katia Gil, Hallman, John, Hehir, Jackie, Heidecke, Johannes, Helyar, Alec, Hu, Haitang, Huet, Romain, Huh, Jacob, Jain, Saachi, Johnson, Zach, Koch, Chris, Kofman, Irina, Kundel, Dominik, Kwon, Jason, Kyrylov, Volodymyr, Le, Elaine Ya, Leclerc, Guillaume, Lennon, James Park, Lessans, Scott, Lezcano-Casado, Mario, Li, Yuanzhi, Li, Zhuohan, Lin, Ji, Liss, Jordan, Lily, null, Liu, null, Liu, Jiancheng, Lu, Kevin, Lu, Chris, Martinovic, Zoran, McCallum, Lindsay, McGrath, Josh, McKinney, Scott, McLaughlin, Aidan, Mei, Song, Mostovoy, Steve, Mu, Tong, Myles, Gideon, Neitz, Alexander, Nichol, Alex, Pachocki, Jakub, Paino, Alex, Palmie, Dana, Pantuliano, Ashley, Parascandolo, Giambattista, Park, Jongsoo, Pathak, Leher, Paz, Carolina, Peran, Ludovic, Pimenov, Dmitry, Pokrass, Michelle, Proehl, Elizabeth, Qiu, Huida, Raila, Gaby, Raso, Filippo, Ren, Hongyu, Richardson, Kimmy, Robinson, David, Rotsted, Bob, Salman, Hadi, Sanjeev, Suvansh, Schwarzer, Max, Sculley, D., Sikchi, Harshit, Simon, Kendal, Singhal, Karan, Song, Yang, Stuckey, Dane, Sun, Zhiqing, Tillet, Philippe, Toizer, Sam, Tsimpourlas, Foivos, Vyas, Nikhil, Wallace, Eric, Wang, Xin, Wang, Miles, Watkins, Olivia, Weil, Kevin, Wendling, Amy, Whinnery, Kevin, Whitney, Cedric, Wong, Hannah, Yang, Lin, Yang, Yu, Yasunaga, Michihiro, Ying, Kristen, Zaremba, Wojciech, Zhan, Wenting, Zhang, Cyril, Zhang, Brian, Zhang, Eddie, Zhao, Shengjia
We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert transformer architecture and are trained using large-scale distillation and reinforcement learning. We optimize the models to have strong agentic capabilities (deep research browsing, python tool use, and support for developer-provided functions), all while using a rendered chat format that enables clear instruction following and role delineation. Both models achieve strong results on benchmarks ranging from mathematics, coding, and safety. We release the model weights, inference implementations, tool environments, and tokenizers under an Apache 2.0 license to enable broad use and further research.
Human-AI collaboration or obedient and often clueless AI in instruct, serve, repeat dynamics?
Saqr, Mohammed, Misiejuk, Kamila, Lรณpez-Pernas, Sonsoles
While research on human-AI collaboration exists, it mainly examined language learning and used traditional counting methods with little attention to evolution and dynamics of collaboration on cognitively demanding tasks. This study examines human-AI interactions while solving a complex problem. Student-AI interactions were qualitatively coded and analyzed with transition network analysis, sequence analysis and partial correlation networks as well as comparison of frequencies using chi-square and Person-residual shaded Mosaic plots to map interaction patterns, their evolution, and their relationship to problem complexity and student performance. Findings reveal a dominant Instructive pattern with interactions characterized by iterative ordering rather than collaborative negotiation. Oftentimes, students engaged in long threads that showed misalignment between their prompts and AI output that exemplified a lack of synergy that challenges the prevailing assumptions about LLMs as collaborative partners. We also found no significant correlations between assignment complexity, prompt length, and student grades suggesting a lack of cognitive depth, or effect of problem difficulty. Our study indicates that the current LLMs, optimized for instruction-following rather than cognitive partnership, compound their capability to act as cognitively stimulating or aligned collaborators. Implications for designing AI systems that prioritize cognitive alignment and collaboration are discussed.
PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins
Chen, Sihan, Lalor, John P., Yang, Yi, Abbasi, Ahmed
While large language models (LLMs) afford new possibilities for user modeling and approximation of human behaviors, they often fail to capture the multidimensional nuances of individual users. In this work, we introduce PersonaTwin, a multi-tier prompt conditioning framework that builds adaptive digital twins by integrating demographic, behavioral, and psychometric data. Using a comprehensive data set in the healthcare context of more than 8,500 individuals, we systematically benchmark PersonaTwin against standard LLM outputs, and our rigorous evaluation unites state-of-the-art text similarity metrics with dedicated demographic parity assessments, ensuring that generated responses remain accurate and unbiased. Experimental results show that our framework produces simulation fidelity on par with oracle settings. Moreover, downstream models trained on persona-twins approximate models trained on individuals in terms of prediction and fairness metrics across both GPT-4o-based and Llama-based models. Together, these findings underscore the potential for LLM digital twin-based approaches in producing realistic and emotionally nuanced user simulations, offering a powerful tool for personalized digital user modeling and behavior analysis.
Constrained Decoding of Diffusion LLMs with Context-Free Grammars
Mรผndler, Niels, Dekoninck, Jasper, Vechev, Martin
Large language models (LLMs) have shown promising performance across diverse domains. Many practical applications of LLMs, such as code completion and structured data extraction, require adherence to syntactic constraints specified by a formal language. Yet, due to their probabilistic nature, LLM output is not guaranteed to adhere to such formal languages. Prior work has proposed constrained decoding as a means to restrict LLM generation to particular formal languages. However, existing works are not applicable to the emerging paradigm of diffusion LLMs, when used in practical scenarios such as the generation of formally correct C++ or JSON output. In this paper we address this challenge and present the first constrained decoding method for diffusion models, one that can handle formal languages captured by context-free grammars. We begin by reducing constrained decoding to the more general additive infilling problem, which asks whether a partial output can be completed to a valid word in the target language. This problem also naturally subsumes the previously unaddressed multi-region infilling constrained decoding. We then reduce this problem to the task of deciding whether the intersection of the target language and a regular language is empty and present an efficient algorithm to solve it for context-free languages. Empirical results on various applications, such as C++ code infilling and structured data extraction in JSON, demonstrate that our method achieves near-perfect syntactic correctness while consistently preserving or improving functional correctness. Importantly, our efficiency optimizations ensure that the computational overhead remains practical.
Training-Free Multimodal Large Language Model Orchestration
Xie, Tianyu, Wu, Yuhang, Luo, Yongdong, Ji, Jiayi, Zheng, Xiawu
Different Multimodal Large Language Models (MLLMs) cannot be integrated into a unified multimodal input-output system directly. In previous work, training has been considered as an inevitable component due to challenges in modal alignment, Text-to-Speech efficiency and other integration issues. In this paper, we introduce Multimodal Large Language Model Orchestration, an effective approach for creating interactive multimodal AI systems without additional training. MLLM Orchestration leverages the inherent reasoning capabilities of large language models to coordinate specialized models through explicit workflows, enabling natural multimodal interactions while maintaining modularity, improving interpretability, and significantly enhancing computational efficiency. Our orchestration framework is built upon three key innovations: (1) a central controller LLM that analyzes user inputs and dynamically routes tasks to appropriate specialized models through carefully designed agents; (2) a parallel Text-to-Speech architecture that enables true full-duplex interaction with seamless interruption handling and natural conversational flow; and (3) a cross-modal memory integration system that maintains coherent context across modalities through intelligent information synthesis and retrieval, selectively avoiding unnecessary modality calls in certain scenarios to improve response speed. Extensive evaluations demonstrate that MLLM Orchestration achieves comprehensive multimodal capabilities without additional training, performance improvements of up to 7.8% over traditional jointly-trained approaches on standard benchmarks, reduced latency by 10.3%, and significantly enhanced interpretability through explicit orchestration processes.
Bridging AI Innovation and Healthcare Needs: Lessons Learned from Incorporating Modern NLP at The BC Cancer Registry
Gondara, Lovedeep, Arbour, Gregory, Ng, Raymond, Simkin, Jonathan, Devji, Shebnum
Automating data extraction from clinical documents offers significant potential to improve efficiency in healthcare settings, yet deploying Natural Language Processing (NLP) solutions presents practical challenges. Drawing upon our experience implementing various NLP models for information extraction and classification tasks at the British Columbia Cancer Registry (BCCR), this paper shares key lessons learned throughout the project lifecycle. We emphasize the critical importance of defining problems based on clear business objectives rather than solely technical accuracy, adopting an iterative approach to development, and fostering deep interdisciplinary collaboration and co-design involving domain experts, end-users, and ML specialists from inception. Further insights highlight the need for pragmatic model selection (including hybrid approaches and simpler methods where appropriate), rigorous attention to data quality (representativeness, drift, annotation), robust error mitigation strategies involving human-in-the-loop validation and ongoing audits, and building organizational AI literacy. These practical considerations, generalizable beyond cancer registries, provide guidance for healthcare organizations seeking to successfully implement AI/NLP solutions to enhance data management processes and ultimately improve patient care and public health outcomes.
Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillation
Ma, Ziyang, Yuan, Qingyue, Zhang, Linhai, Zhou, Deyu
Previous chain-of-thought (CoT) distillation methods primarily focused on enhancing the reasoning capabilities of Small Language Models (SLMs) by utilizing high-quality rationales generated by powerful Large Language Models (LLMs, e.g., GPT-4). However, few works have noted the negative effects on SLM safety brought by the training, which are revealed in this study. Although there are works on safety alignment that fine-tune language models or manipulate model weights to defend against harmful inputs, they require extra computation or annotated data, and probably impact the reasoning ability of SLMs. In this paper, we investigate how to maintain the safety of SLMs during the CoT distillation process. Specifically, we propose a safe distillation method, Slow Tuning and Low-Entropy Masking Distillation (SLowED), containing two modules: Slow Tuning and Low-Entropy Masking. Slow Tuning scales down the magnitude of model weight changes to optimize the model weights in the neighboring space near the initial weight distribution. Low-Entropy Masking masks low-entropy tokens, which are regarded as unnecessary learning targets, to exclude them from fine-tuning. Experiments on three SLMs (Qwen2.5-1.5B, Llama-3.2-1B, BLOOM-1.1B) across reasoning benchmarks (BBH, BB-Sub, ARC, AGIEval) and safety evaluation (AdvBench) show that SLowED retains the safety of SLMs and comparably improves their reasoning capability compared to existing distillation methods. Furthermore, our ablation study presents the effectiveness of Slow Tuning and Low-Entropy Masking, with the former maintaining the model's safety in the early stage and the latter prolonging the safe training epochs.
TimeMKG: Knowledge-Infused Causal Reasoning for Multivariate Time Series Modeling
Sun, Yifei, Liu, Junming, Chen, Yirong, Yan, Xuefeng, Wang, Ding
Multivariate time series data typically comprises two distinct modalities: variable semantics and sampled numerical observations. Traditional time series models treat variables as anonymous statistical signals, overlooking the rich semantic information embedded in variable names and data descriptions. However, these textual descriptors often encode critical domain knowledge that is essential for robust and interpretable modeling. Here we present TimeMKG, a mul-timodal causal reasoning framework that elevates time series modeling from low-level signal processing to knowledge informed inference. TimeMKG employs large language models to interpret variable semantics and constructs structured Multivariate Knowledge Graphs that capture inter-variable relationships. A dual-modality encoder separately models the semantic prompts--generated from knowledge graph triplets--and the statistical patterns from historical time series. Cross-modality attention aligns and fuses these representations at the variable level, injecting causal priors into downstream tasks such as forecasting and classification--providing explicit and interpretable priors to guide model reasoning. The experiment in diverse datasets demonstrates that incorporating variable-level knowledge significantly improves both predictive performance and generalization.
E3-Rewrite: Learning to Rewrite SQL for Executability, Equivalence,and Efficiency
Xu, Dongjie, Cui, Yue, Shi, Weijie, Ma, Qingzhi, Guo, Hanghui, Li, Jiaming, Zhao, Yao, Zhang, Ruiyuan, Di, Shimin, Zhu, Jia, Zheng, Kai, Xu, Jiajie
SQL query rewriting aims to reformulate a query into a more efficient form while preserving equivalence. Most existing methods rely on predefined rewrite rules. However, such rule-based approaches face fundamental limitations: (1) fixed rule sets generalize poorly to novel query patterns and struggle with complex queries; (2) a wide range of effective rewriting strategies cannot be fully captured by declarative rules. To overcome these issues, we propose using large language models (LLMs) to generate rewrites. LLMs can capture complex strategies, such as evaluation reordering and CTE rewriting. Despite this potential, directly applying LLMs often results in performance regressions or non-equivalent rewrites due to a lack of execution awareness and semantic grounding. To address these challenges, We present E3-Rewrite, an LLM-based SQL rewriting framework that produces executable, equivalent, and efficient queries. It integrates two core components: a context construction module and a reinforcement learning framework. First, the context module leverages execution plans and retrieved demonstrations to build bottleneck-aware prompts that guide inference-time rewriting. Second, we design a reward function targeting executability, equivalence, and efficiency, evaluated via syntax checks, equivalence verification, and cost estimation. Third, to ensure stable multi-objective learning, we adopt a staged curriculum that first emphasizes executability and equivalence, then gradually incorporates efficiency. Across multiple SQL benchmarks, our experiments demonstrate that E3-Rewrite can shorten query execution time by as much as 25.6% relative to leading baselines, while also producing up to 24.4% more rewrites that meet strict equivalence criteria. These gains extend to challenging query patterns that prior approaches could not effectively optimize.