Deep Learning
Hermes 4 Technical Report
Teknium, Ryan, Jin, Roger, Suphavadeeprasit, Jai, Mahan, Dakota, Quesnelle, Jeffrey, Li, Joe, Guang, Chen, Sands, Shannon, Malhotra, Karan
We present Hermes 4, a family of hybrid reasoning models that combine structured, multi-turn reasoning with broad instruction-following ability. We describe the challenges encountered during data curation, synthesis, training, and evaluation, and outline the solutions employed to address these challenges at scale. We comprehensively evaluate across mathematical reasoning, coding, knowledge, comprehension, and alignment benchmarks, and we report both quantitative performance and qualitative behavioral analysis. To support open research, all model weights are published publicly at https://huggingface.co/collections/NousResearch/hermes-4-collection-68a731bfd452e20816725728
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
NVIDIA, null, :, null, Basant, Aarti, Khairnar, Abhijit, Paithankar, Abhijit, Khattar, Abhinav, Renduchintala, Adithya, Malte, Aditya, Bercovich, Akhiad, Hazare, Akshay, Rico, Alejandra, Ficek, Aleksander, Kondratenko, Alex, Shaposhnikov, Alex, Bukharin, Alexander, Taghibakhshi, Ali, Barton, Amelia, Mahabaleshwarkar, Ameya Sunil, Shen, Amy, Tao, Andrew, Guan, Ann, Shors, Anna, Mandarwal, Anubhav, Mehta, Arham, Venkatesan, Arun, Sharabiani, Ashton, Aithal, Ashwath, Poojary, Ashwin, Dattagupta, Ayush, Buddharaju, Balaram, Zhu, Banghua, Simkin, Barnaby, Kartal, Bilal, Rouhani, Bita Darvish, Chen, Bobby, Ginsburg, Boris, Norick, Brandon, Yu, Brian, Catanzaro, Bryan, Wang, Charles, Truong, Charlie, Mungekar, Chetan, Patel, Chintan, Alexiuk, Chris, Munley, Christian, Parisien, Christopher, Su, Dan, Afrimi, Daniel, Korzekwa, Daniel, Rohrer, Daniel, Gitman, Daria, Mosallanezhad, David, Narayanan, Deepak, Rekesh, Dima, Yared, Dina, Pykhtar, Dmytro, Ahn, Dong, Riach, Duncan, Long, Eileen, Ning, Elliott, Chung, Eric, Galinkin, Erick, Bakhturina, Evelina, Prasad, Gargi, Shen, Gerald, Qian, Haifeng, Elisha, Haim, Sharma, Harsh, Ross, Hayley, Ngo, Helen, Sahota, Herman, Wang, Hexin, Shin, Hoo Chang, Huang, Hua, Cunningham, Iain, Gitman, Igor, Moshkov, Ivan, Jung, Jaehun, Kautz, Jan, Scowcroft, Jane Polak, Casper, Jared, Zhang, Jian, Zeng, Jiaqi, Zhang, Jimmy, Xue, Jinze, Huang, Jocelyn, Conway, Joey, Kamalu, John, Cohen, Jonathan, Jennings, Joseph, Vialard, Julien Veron, Yi, Junkeun, Parmar, Jupinder, Briski, Kari, Cheung, Katherine, Luna, Katherine, Wyss, Keith, Santhanam, Keshav, Kong, Kezhi, Pawelec, Krzysztof, Anik, Kumar, Li, Kunlun, Ahmadian, Kushan, McAfee, Lawrence, Sleiman, Laya, Derczynski, Leon, Vega, Luis, de Melo, Maer Rodrigues, Sreedhar, Makesh Narsimhan, Chochowski, Marcin, Cai, Mark, Kliegl, Markus, Stepniewska-Dziubinska, Marta, Novikov, Matvei, Samadi, Mehrzad, Price, Meredith, Boubdir, Meriem, Boone, Michael, Evans, Michael, Bien, Michal, Zawalski, Michal, Martinez, Miguel, Chrzanowski, Mike, Shoeybi, Mohammad, Patwary, Mostofa, Dhameja, Namit, Assaf, Nave, Habibi, Negar, Bhatia, Nidhi, Pope, Nikki, Tajbakhsh, Nima, Juluru, Nirmal Kumar, Rybakov, Oleg, Hrinchuk, Oleksii, Kuchaiev, Oleksii, Olabiyi, Oluwatobi, Ribalta, Pablo, Subramanian, Padmavathy, Chadha, Parth, Molchanov, Pavlo, Dykas, Peter, Jin, Peter, Bialecki, Piotr, Januszewski, Piotr, Thalasta, Pradeep, Gaikwad, Prashant, Varshney, Prasoon, Gundecha, Pritam, Tredak, Przemek, Mahabadi, Rabeeh Karimi, Patel, Rajen, El-Yaniv, Ran, Rajan, Ranjit, Cheruvu, Ria, Shahbazyan, Rima, Borkar, Ritika, Gala, Ritu, Waleffe, Roger, Zhang, Ruoxi, Hewett, Russell J., Prenger, Ryan, Jain, Sahil, Kriman, Samuel, Satheesh, Sanjeev, Kaji, Saori, Yurick, Sarah, Muralidharan, Saurav, Narenthiran, Sean, Bak, Seonmyeong, Sameni, Sepehr, Han, Seungju, Ramasamy, Shanmugam, Ghosh, Shaona, Sreenivas, Sharath Turuvekere, Thomas, Shelby, Diao, Shizhe, Gopal, Shreya, Prabhumoye, Shrimai, Toshniwal, Shubham, Ding, Shuoyang, Singh, Siddharth, Jain, Siddhartha, Majumdar, Somshubra, Singhal, Soumye, Alborghetti, Stefania, Akter, Syeda Nahida, Kong, Terry, Moon, Tim, Hliwiak, Tomasz, Asida, Tomer, Wang, Tony, Konuk, Tugrul, Vashishth, Twinkle, Poon, Tyler, Karpas, Udi, Noroozi, Vahid, Srinivasan, Venkat, Korthikanti, Vijay, Fugro, Vikram, Kalluru, Vineeth, Kurin, Vitaly, Lavrukhin, Vitaly, Ahmad, Wasi Uddin, Du, Wei, Byeon, Wonmin, Lu, Ximing, Dong, Xin, Karnati, Yashaswi, Choi, Yejin, Zhang, Yian, Lin, Ying, Fu, Yonggan, Suhara, Yoshi, Dong, Zhen, Li, Zhiyu, Zhu, Zhongbo, Chen, Zijia
We introduce Nemotron-Nano-9B-v2, a hybrid Mamba-Transformer language model designed to increase throughput for reasoning workloads while achieving state-of-the-art accuracy compared to similarly-sized models. Nemotron-Nano-9B-v2 builds on the Nemotron-H architecture, in which the majority of the self-attention layers in the common Transformer architecture are replaced with Mamba-2 layers, to achieve improved inference speed when generating the long thinking traces needed for reasoning. We create Nemotron-Nano-9B-v2 by first pre-training a 12-billion-parameter model (Nemotron-Nano-12B-v2-Base) on 20 trillion tokens using an FP8 training recipe. After aligning Nemotron-Nano-12B-v2-Base, we employ the Minitron strategy to compress and distill the model with the goal of enabling inference on up to 128k tokens on a single NVIDIA A10G GPU (22GiB of memory, bfloat16 precision). Compared to existing similarly-sized models (e.g., Qwen3-8B), we show that Nemotron-Nano-9B-v2 achieves on-par or better accuracy on reasoning benchmarks while achieving up to 6x higher inference throughput in reasoning settings like 8k input and 16k output tokens. We are releasing Nemotron-Nano-9B-v2, Nemotron-Nano12B-v2-Base, and Nemotron-Nano-9B-v2-Base checkpoints along with the majority of our pre- and post-training datasets on Hugging Face.
SpatialViz-Bench: An MLLM Benchmark for Spatial Visualization
Wang, Siting, Pei, Minnan, Sun, Luoyang, Deng, Cheng, Shao, Kun, Tian, Zheng, Zhang, Haifeng, Wang, Jun
Humans can directly imagine and manipulate visual images in their minds, a capability known as spatial visualization. While multi-modal Large Language Models (MLLMs) support imagination-based reasoning, spatial visualization remains insufficiently evaluated, typically embedded within broader mathematical and logical assessments. Existing evaluations often rely on IQ tests or math competitions that may overlap with training data, compromising assessment reliability. To this end, we introduce SpatialViz-Bench, a comprehensive multi-modal benchmark for spatial visualization with 12 tasks across 4 sub-abilities, comprising 1,180 automatically generated problems. Our evaluation of 33 state-of-the-art MLLMs not only reveals wide performance variations and demonstrates the benchmark's strong discriminative power, but also uncovers counter-intuitive findings: models show difficulty perception misaligned with human intuition, exhibit dramatic 2Dto-3D performance cliffs, default to formulaic derivation over visualization, and paradoxically suffer performance degradation from Chain-of-Thought prompting in open-source models. Through statistical and qualitative analysis of error types, SpatialViz-Bench demonstrates that state-of-the-art MLLMs continue to exhibit deficiencies in spatial visualization tasks, thereby addressing a significant lacuna in the field. The benchmark data and evaluation code are publicly available.
Toward a Robust and Generalizable Metamaterial Foundation Model
Kim, Namjung, Lee, Dongseok, Yu, Jongbin, Cho, Sung Woong, Lee, Dosung, Park, Yesol, Hong, Youngjoon
Advances in material functionalities drive innovations across various fields, where metamaterials-defined by structure rather than composition-are leading the way. Despite the rise of artificial intelligence (AI)-driven design strategies, their impact is limited by task-specific retraining, poor out-of-distribution(OOD) generalization, and the need for separate models for forward and inverse design. To address these limitations, we introduce the Metamaterial Foundation Model (MetaFO), a Bayesian transformer-based foundation model inspired by large language models. MetaFO learns the underlying mechanics of metamaterials, enabling probabilistic, zero-shot predictions across diverse, unseen combinations of material properties and structural responses. It also excels in nonlinear inverse design, even under OOD conditions. By treating metamaterials as an operator that maps material properties to structural responses, MetaFO uncovers intricate structure-property relationships and significantly expands the design space. This scalable and generalizable framework marks a paradigm shift in AI-driven metamaterial discovery, paving the way for next-generation innovations.
TPTT: Transforming Pretrained Transformers into Titans
Transformer-based large language models (LLMs) have achieved strong performance across many natural language processing tasks. Nonetheless, their quadratic computational and memory requirements, particularly in self-attention layers, pose challenges for efficient inference on long contexts and for deployment in resource-limited environments. We present TPTT (Transforming Pretrained Transformers into Titans), a framework designed to augment pretrained Transformers with linearized attention (LiZA) and internal memory gating via Memory as Gate (MaG), applied without full retraining. TPTT supports parameter-efficient fine-tuning (LoRA) and integrates with standard toolkits such as Hugging Face Transformers. We evaluated TPTT on several pretrained models, including Llama-1B, OlMoE-1B-7B, Qwen2.5-1.5B, Gemma3-270m, OpenELM-1.3B, and Mistral-7B, in order to assess applicability across architectures of different scales. Experiments on models with approximately 1 billion parameters, evaluated primarily on the MMLU benchmark, suggest potential improvements in both efficiency and accuracy compared to baseline models. For example, Titans-Llama-1B exhibited up to a 20\% relative increase in Exact Match scores in one-shot evaluation. An additional finding is that it is possible to convert a quadratic-attention model into a purely linear-attention model using the DeltaProduct mechanism. All training runs were carried out with modest computational resources. These preliminary findings indicate that TPTT may help adapt pretrained LLMs for long-context tasks with limited overhead. Further studies on larger models and a broader set of benchmarks will be necessary to evaluate the generality and robustness of the framework. Code is available at https://github.com/fabienfrfr/tptt . Python package at https://pypi.org/project/tptt/ .
EmoPerso: Enhancing Personality Detection with Self-Supervised Emotion-Aware Modelling
Shen, Lingzhi, Cai, Xiaohao, Long, Yunfei, Razzak, Imran, Chen, Guanming, Jameel, Shoaib
Personality detection from text is commonly performed by analysing users' social media posts. However, existing methods heavily rely on large-scale annotated datasets, making it challenging to obtain high-quality personality labels. Moreover, most studies treat emotion and personality as independent variables, overlooking their interactions. In this paper, we propose a novel self-supervised framework, EmoPerso, which improves personality detection through emotion-aware modelling. EmoPerso first leverages generative mechanisms for synthetic data augmentation and rich representation learning. It then extracts pseudo-labeled emotion features and jointly optimizes them with personality prediction via multi-task learning. A cross-attention module is employed to capture fine-grained interactions between personality traits and the inferred emotional representations. To further refine relational reasoning, EmoPerso adopts a self-taught strategy to enhance the model's reasoning capabilities iteratively. Extensive experiments on two benchmark datasets demonstrate that EmoPerso surpasses state-of-the-art models. The source code is available at https://github.com/slz0925/EmoPerso.
An Ensemble Classification Approach in A Multi-Layered Large Language Model Framework for Disease Prediction
Hamdi, Ali, Mohamed, Malak, Emad, Rokaia, Shaban, Khaled
Social telehealth has made remarkable progress in healthcare by allowing patients to post symptoms and participate in medical consultations remotely. Users frequently post symptoms on social media and online health platforms, creating a huge repository of medical data that can be leveraged for disease classification. Large language models (LLMs) such as LLAMA3 and GPT-3.5, along with transformer-based models like BERT, have demonstrated strong capabilities in processing complex medical text. In this study, we evaluate three Arabic medical text preprocessing methods such as summarization, refinement, and Named Entity Recognition (NER) before applying fine-tuned Arabic transformer models (CAMeLBERT, AraBERT, and AsafayaBERT). To enhance robustness, we adopt a majority voting ensemble that combines predictions from original and preprocessed text representations. This approach achieved the best classification accuracy of 80.56%, thus showing its effectiveness in leveraging various text representations and model predictions to improve the understanding of medical texts. To the best of our knowledge, this is the first work that integrates LLM-based preprocessing with fine-tuned Arabic transformer models and ensemble learning for disease classification in Arabic social telehealth data.
VASSO: Variance Suppression for Sharpness-Aware Minimization
Li, Bingcong, Zhang, Yilang, Giannakis, Georgios B.
Sharpness-aware minimization (SAM) has well-documented merits in enhancing generalization of deep neural network models. Accounting for sharpness in the loss function geometry, where neighborhoods of `flat minima' heighten generalization ability, SAM seeks `flat valleys' by minimizing the maximum loss provoked by an adversarial perturbation within the neighborhood. Although critical to account for sharpness of the loss function, in practice SAM suffers from `over-friendly adversaries,' which can curtail the outmost level of generalization. To avoid such `friendliness,' the present contribution fosters stabilization of adversaries through variance suppression (VASSO). VASSO offers a general approach to provably stabilize adversaries. In particular, when integrating VASSO with SAM, improved generalizability is numerically validated on extensive vision and language tasks. Once applied on top of a computationally efficient SAM variant, VASSO offers a desirable generalization-computation tradeoff.
From Noisy Labels to Intrinsic Structure: A Geometric-Structural Dual-Guided Framework for Noise-Robust Medical Image Segmentation
Wang, Tao, Zhang, Zhenxuan, Zhou, Yuanbo, Zhang, Xinlin, Chen, Yuanbin, Tan, Tao, Yang, Guang, Tong, Tong
The effectiveness of convolutional neural networks in medical image segmentation relies on large-scale, high-quality annotations, which are costly and time-consuming to obtain. Even expert-labeled datasets inevitably contain noise arising from subjectivity and coarse delineations, which disrupt feature learning and adversely impact model performance. To address these challenges, this study propose a Geometric-Structural Dual-Guided Network (GSD-Net), which integrates geometric and structural cues to improve robustness against noisy annotations. It incorporates a Geometric Distance-Aware module that dynamically adjusts pixel-level weights using geometric features, thereby strengthening supervision in reliable regions while suppressing noise. A Structure-Guided Label Refinement module further refines labels with structural priors, and a Knowledge Transfer module enriches supervision and improves sensitivity to local details. To comprehensively assess its effectiveness, we evaluated GSD-Net on six publicly available datasets: four containing three types of simulated label noise, and two with multi-expert annotations that reflect real-world subjectivity and labeling inconsistencies. Experimental results demonstrate that GSD-Net achieves state-of-the-art performance under noisy annotations, achieving improvements of 2.52% on Kvasir, 22.76% on Shenzhen, 8.87% on BU-SUC, and 4.59% on BraTS2020 under SR simulated noise. The codes of this study are available at https://github.com/ortonwang/GSD-Net.
A Survey: Towards Privacy and Security in Mobile Large Language Models
Xu, Honghui, Li, Kaiyang, Chen, Wei, Zheng, Danyang, Li, Zhiyuan, Cai, Zhipeng
--Mobile Large Language Models (LLMs) are revolutionizing diverse fields such as healthcare, finance, and education with their ability to perform advanced natural language processing tasks on-the-go. However, the deployment of these models in mobile and edge environments introduces significant challenges related to privacy and security due to their resource-intensive nature and the sensitivity of the data they process. This survey provides a comprehensive overview of privacy and security issues associated with mobile LLMs, systematically categorizing existing solutions such as differential privacy, federated learning, and prompt encryption. Furthermore, we analyze vulnerabilities unique to mobile LLMs, including adversarial attacks, membership inference, and side-channel attacks, offering an in-depth comparison of their effectiveness and limitations. T o bridge this gap, we propose potential applications, discuss open challenges, and suggest future research directions, paving the way for the development of trustworthy, privacy-compliant, and scalable mobile LLM systems. The advent of mobile Large Language Models (LLMs) represents a significant milestone in the evolution of AI, enabling advanced natural language processing capabilities to be deployed in mobile and edge environments [1]-[3]. By bringing powerful AI tools closer to end-users, mobile LLMs are revolutionizing industries such as healthcare [4], finance [5], and education [6] with real-time, on-device applications.