Deep Learning
Towards Efficient Prompt-based Continual Learning in Distributed Medical AI
Modern AI models achieve state-of-the-art performance with large-scale, high-quality datasets; however, ethical, social, and institutional constraints in the medical domain severely restrict data sharing, rendering centralized learning nearly impossible. Each institution must incrementally update models using only local data. Traditional training overfits new samples and suffers from catastrophic forgetting, losing previously acquired knowledge. Medical data distributions also shift due to varying diagnostic equipment and demographics. Although continual learning (CL) has advanced, most methods address natural images, leaving medical-domain-specific CL underexplored. We propose a prompt-based continual learning (PCL) approach featuring a unified prompt pool with a minimal expansion strategy: by expanding and freezing a subset of prompts, our method reduces computational overhead, and a novel regularization term balances retention and adaptation. Experiments on three diabetic retinopathy datasets Aptos2019, LI2019, and Diabetic Retinopathy Detection show our model improves final classification accuracy by at least 10% and F1-score by 9 points over state-of-the-art approaches while lowering inference cost. We anticipate this study will drive sustainable medical AI advances, enabling real-time diagnosis, patient monitoring, and telemedicine applications in distributed healthcare. Code will be released upon acceptance
Apriel-Nemotron-15B-Thinker
Radhakrishna, Shruthan, Parikh, Soham, Sarda, Gopal, Turkkan, Anil, Vohra, Quaizar, Li, Raymond, Jhamb, Dhruv, Ogueji, Kelechi, Shukla, Aanjaneya, Bamgbose, Oluwanifemi, Liang, Toby, Kumar, Luke, Ostapenko, Oleksiy, Malay, Shiva Krishna Reddy, Tiwari, Aman, Bogavelli, Tara, Yadav, Vikas, Mehta, Jash, Mittal, Saloni, Kalkunte, Akshay, Pattnaik, Pulkit, Slimi, Khalil, Sreeram, Anirudh, Nair, Jishnu, Oladipo, Akintunde, Maiya, Shashank, Mahajan, Khyati, Maheshwary, Rishabh, Hashemi, Masoud, Mudumba, Sai Rajeswar, Madhusudhan, Sathwik Tejaswi, Scholak, Torsten, Paquet, Sebastien, Davasam, Sagar, Sunkara, Srinivas
While large language models (LLMs) have achieved remarkable reasoning capabilities across domains like code, math and other enterprise tasks, their significant memory and computational costs often preclude their use in practical enterprise settings. To this end, we introduce Apriel-Nemotron-15B-Thinker, a 15-billion parameter model in the ServiceNow Apriel SLM series that achieves performance against medium sized state-of-the-art models such as o1-mini, QWQ32B, and EXAONE-Deep-32B while maintaining only half the memory footprint of those alternatives. Apriel-Nemotron-15B-Thinker model is trained in a four stage training pipeline including 1) Base Model upscaling, 2) Continual Pre-training 3) Supervised Fine-tuning (SFT) and 4) Reinforcement Learning using GRPO. Comprehensive evaluations across a diverse suite of benchmarks consistently demonstrate that our Apriel-Nemotron-15B-Thinker model matches or exceeds the performance of its 32-billion parameter counterparts, despite being less than half their size.
iWatchRoad: Scalable Detection and Geospatial Visualization of Potholes for Smart Cities
Sahoo, Rishi Raj, Mohanty, Surbhi Saswati, Mishra, Subhankar
Potholes on the roads are a serious hazard and maintenance burden. This poses a significant threat to road safety and vehicle longevity, especially on the diverse and under-maintained roads of India. In this paper, we present a complete end-to-end system called iWatchRoad for automated pothole detection, Global Positioning System (GPS) tagging, and real time mapping using OpenStreetMap (OSM). We curated a large, self-annotated dataset of over 7,000 frames captured across various road types, lighting conditions, and weather scenarios unique to Indian environments, leveraging dashcam footage. This dataset is used to fine-tune, Ultralytics You Only Look Once (YOLO) model to perform real time pothole detection, while a custom Optical Character Recognition (OCR) module was employed to extract timestamps directly from video frames. The timestamps are synchronized with GPS logs to geotag each detected potholes accurately. The processed data includes the potholes' details and frames as metadata is stored in a database and visualized via a user friendly web interface using OSM. iWatchRoad not only improves detection accuracy under challenging conditions but also provides government compatible outputs for road assessment and maintenance planning through the metadata visible on the website. Our solution is cost effective, hardware efficient, and scalable, offering a practical tool for urban and rural road management in developing regions, making the system automated. iWatchRoad is available at https://smlab.niser.ac.in/project/iwatchroad
CleanCTG: A Deep Learning Model for Multi-Artefact Detection and Reconstruction in Cardiotocography
Wong, Sheng, Albert, Beth, Jones, Gabriel Davis
Cardiotocography (CTG) is essential for fetal monitoring but is frequently compromised by diverse artefacts which obscure true fetal heart rate (FHR) patterns and can lead to misdiagnosis or delayed intervention. Current deep-learning approaches typically bypass comprehensive noise handling, applying minimal preprocessing or focusing solely on downstream classification, while traditional methods rely on simple interpolation or rule-based filtering that addresses only missing samples and fail to correct complex artefact types. We present CleanCTG, an end-to-end dual-stage model that first identifies multiple artefact types via multi-scale convolution and context-aware cross-attention, then reconstructs corrupted segments through artefact-specific correction branches. Training utilised over 800,000 minutes of physiologically realistic, synthetically corrupted CTGs derived from expert-verified "clean" recordings. On synthetic data, CleanCTG achieved perfect artefact detection (AU-ROC = 1.00) and reduced mean squared error (MSE) on corrupted segments to 2.74 x 10^-4 (clean-segment MSE = 2.40 x 10^-6), outperforming the next best method by more than 60%. External validation on 10,190 minutes of clinician-annotated segments yielded AU-ROC = 0.95 (sensitivity = 83.44%, specificity 94.22%), surpassing six comparator classifiers. Finally, when integrated with the Dawes-Redman system on 933 clinical CTG recordings, denoised traces increased specificity (from 80.70% to 82.70%) and shortened median time to decision by 33%. These findings suggest that explicit artefact removal and signal reconstruction can both maintain diagnostic accuracy and enable shorter monitoring sessions, offering a practical route to more reliable CTG interpretation.
gpt-oss-120b & gpt-oss-20b Model Card
OpenAI, null, :, null, Agarwal, Sandhini, Ahmad, Lama, Ai, Jason, Altman, Sam, Applebaum, Andy, Arbus, Edwin, Arora, Rahul K., Bai, Yu, Baker, Bowen, Bao, Haiming, Barak, Boaz, Bennett, Ally, Bertao, Tyler, Brett, Nivedita, Brevdo, Eugene, Brockman, Greg, Bubeck, Sebastien, Chang, Che, Chen, Kai, Chen, Mark, Cheung, Enoch, Clark, Aidan, Cook, Dan, Dukhan, Marat, Dvorak, Casey, Fives, Kevin, Fomenko, Vlad, Garipov, Timur, Georgiev, Kristian, Glaese, Mia, Gogineni, Tarun, Goucher, Adam, Gross, Lukas, Guzman, Katia Gil, Hallman, John, Hehir, Jackie, Heidecke, Johannes, Helyar, Alec, Hu, Haitang, Huet, Romain, Huh, Jacob, Jain, Saachi, Johnson, Zach, Koch, Chris, Kofman, Irina, Kundel, Dominik, Kwon, Jason, Kyrylov, Volodymyr, Le, Elaine Ya, Leclerc, Guillaume, Lennon, James Park, Lessans, Scott, Lezcano-Casado, Mario, Li, Yuanzhi, Li, Zhuohan, Lin, Ji, Liss, Jordan, Lily, null, Liu, null, Liu, Jiancheng, Lu, Kevin, Lu, Chris, Martinovic, Zoran, McCallum, Lindsay, McGrath, Josh, McKinney, Scott, McLaughlin, Aidan, Mei, Song, Mostovoy, Steve, Mu, Tong, Myles, Gideon, Neitz, Alexander, Nichol, Alex, Pachocki, Jakub, Paino, Alex, Palmie, Dana, Pantuliano, Ashley, Parascandolo, Giambattista, Park, Jongsoo, Pathak, Leher, Paz, Carolina, Peran, Ludovic, Pimenov, Dmitry, Pokrass, Michelle, Proehl, Elizabeth, Qiu, Huida, Raila, Gaby, Raso, Filippo, Ren, Hongyu, Richardson, Kimmy, Robinson, David, Rotsted, Bob, Salman, Hadi, Sanjeev, Suvansh, Schwarzer, Max, Sculley, D., Sikchi, Harshit, Simon, Kendal, Singhal, Karan, Song, Yang, Stuckey, Dane, Sun, Zhiqing, Tillet, Philippe, Toizer, Sam, Tsimpourlas, Foivos, Vyas, Nikhil, Wallace, Eric, Wang, Xin, Wang, Miles, Watkins, Olivia, Weil, Kevin, Wendling, Amy, Whinnery, Kevin, Whitney, Cedric, Wong, Hannah, Yang, Lin, Yang, Yu, Yasunaga, Michihiro, Ying, Kristen, Zaremba, Wojciech, Zhan, Wenting, Zhang, Cyril, Zhang, Brian, Zhang, Eddie, Zhao, Shengjia
We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert transformer architecture and are trained using large-scale distillation and reinforcement learning. We optimize the models to have strong agentic capabilities (deep research browsing, python tool use, and support for developer-provided functions), all while using a rendered chat format that enables clear instruction following and role delineation. Both models achieve strong results on benchmarks ranging from mathematics, coding, and safety. We release the model weights, inference implementations, tool environments, and tokenizers under an Apache 2.0 license to enable broad use and further research.
Uncovering Latent Connections in Indigenous Heritage: Semantic Pipelines for Cultural Preservation in Brazil
Zerkowski, Luis Vitor, Hirata, Nina S. T.
Indigenous communities face ongoing challenges in preserving their cultural heritage, particularly in the face of systemic marginalization and urban development. In Brazil, the Museu Nacional dos Povos Indigenas through the Tainacan platform hosts the country's largest online collection of Indigenous objects and iconographies, providing a critical resource for cultural engagement. Using publicly available data from this repository, we present a data-driven initiative that applies artificial intelligence to enhance accessibility, interpretation, and exploration. We develop two semantic pipelines: a visual pipeline that models image-based similarity and a textual pipeline that captures semantic relationships from item descriptions. These embedding spaces are projected into two dimensions and integrated into an interactive visualization tool we also developed. In addition to similarity-based navigation, users can explore the collection through temporal and geographic lenses, enabling both semantic and contextualized perspectives. The system supports curatorial tasks, aids public engagement, and reveals latent connections within the collection. This work demonstrates how AI can ethically contribute to cultural preservation practices.
PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins
Chen, Sihan, Lalor, John P., Yang, Yi, Abbasi, Ahmed
While large language models (LLMs) afford new possibilities for user modeling and approximation of human behaviors, they often fail to capture the multidimensional nuances of individual users. In this work, we introduce PersonaTwin, a multi-tier prompt conditioning framework that builds adaptive digital twins by integrating demographic, behavioral, and psychometric data. Using a comprehensive data set in the healthcare context of more than 8,500 individuals, we systematically benchmark PersonaTwin against standard LLM outputs, and our rigorous evaluation unites state-of-the-art text similarity metrics with dedicated demographic parity assessments, ensuring that generated responses remain accurate and unbiased. Experimental results show that our framework produces simulation fidelity on par with oracle settings. Moreover, downstream models trained on persona-twins approximate models trained on individuals in terms of prediction and fairness metrics across both GPT-4o-based and Llama-based models. Together, these findings underscore the potential for LLM digital twin-based approaches in producing realistic and emotionally nuanced user simulations, offering a powerful tool for personalized digital user modeling and behavior analysis.
PTQAT: A Hybrid Parameter-Efficient Quantization Algorithm for 3D Perception Tasks
Wang, Xinhao, Lin, Zhiwei, Xia, Zhongyu, Wang, Yongtao
Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT) represent two mainstream model quantization approaches. However, PTQ often leads to unacceptable performance degradation in quantized models, while QAT imposes substantial GPU memory requirements and extended training time due to weight fine-tuning. In this paper, we propose PTQAT, a novel general hybrid quantization algorithm for the efficient deployment of 3D perception networks. T o address the speed-accuracy trade-off between PTQ and QAT, our method selects critical layers for QAT fine-tuning and performs PTQ on the remaining layers. Contrary to intuition, fine-tuning the layers with smaller output discrepancies before and after quantization, rather than those with larger discrepancies, actually leads to greater improvements in the model's quantization accuracy. This means we better compensate for quantization errors during their propagation, rather than addressing them at the point where they occur . The proposed PTQAT achieves similar performance to QAT with more efficiency by freezing nearly 50% of quantifiable layers. Additionally, PTQAT is a universal quantization method that supports various quantization bit widths (4 bits) as well as different model architectures, including CNNs and Transformers. The experimental results on nuScenes across diverse 3D perception tasks, including object detection, semantic segmentation, and occupancy prediction, show that our method consistently outperforms QAT-only baselines.
Driving Accurate Allergen Prediction with Protein Language Models and Generalization-Focused Evaluation
Wong, Brian Shing-Hei, Kim, Joshua Mincheol, Fung, Sin-Hang, Xiong, Qing, Ao, Kelvin Fu-Kiu, Wei, Junkang, Wang, Ran, Wang, Dan Michelle, Zhou, Jingying, Feng, Bo, Cheng, Alfred Sze-Lok, Yip, Kevin Y., Tsui, Stephen Kwok-Wing, Cao, Qin
Allergens, typically proteins capable of triggering adverse immune responses, represent a significant public health challenge. To accurately identify allergen proteins, we introduce Applm (Allergen Prediction with Protein Language Models), a computational framework that leverages the 100-billion parameter xTrimoPGLM protein language model. We show that Applm consistently outperforms seven state-of-the-art methods in a diverse set of tasks that closely resemble difficult real-world scenarios. These include identifying novel allergens that lack similar examples in the training set, differentiating between allergens and non-allergens among homologs with high sequence similarity, and assessing functional consequences of mutations that create few changes to the protein sequences. Our analysis confirms that xTrimoPGLM, originally trained on one trillion tokens to capture general protein sequence characteristics, is crucial for Applm's performance by detecting important differences among protein sequences. In addition to providing Applm as open-source software, we also provide our carefully curated benchmark datasets to facilitate future research.
When Explainability Meets Privacy: An Investigation at the Intersection of Post-hoc Explainability and Differential Privacy in the Context of Natural Language Processing
Dhaini, Mahdi, Meisenbacher, Stephen, Erdogan, Ege, Matthes, Florian, Kasneci, Gjergji
In the study of trustworthy Natural Language Processing (NLP), a number of important research fields have emerged, including that of explainability and privacy. While research interest in both explainable and privacy-preserving NLP has increased considerably in recent years, there remains a lack of investigation at the intersection of the two. This leaves a considerable gap in understanding of whether achieving both explainability and privacy is possible, or whether the two are at odds with each other. In this work, we conduct an empirical investigation into the privacy-explainability trade-off in the context of NLP, guided by the popular overarching methods of Differential Privacy (DP) and Post-hoc Explainability. Our findings include a view into the intricate relationship between privacy and explainability, which is formed by a number of factors, including the nature of the downstream task and choice of the text privatization and explainability method. In this, we highlight the potential for privacy and explainability to co-exist, and we summarize our findings in a collection of practical recommendations for future work at this important intersection.