Statistical Learning
Occam's model: Selecting simpler representations for better transferability estimation
Singh, Prabhant, Hess, Sibylle, Vanschoren, Joaquin
Fine-tuning models that have been pre-trained on large datasets has become a cornerstone of modern machine learning workflows. With the widespread availability of online model repositories, such as Hugging Face, it is now easier than ever to fine-tune pre-trained models for specific tasks. This raises a critical question: which pre-trained model is most suitable for a given task? This problem is called transferability estimation. In this work, we introduce two novel and effective metrics for estimating the transferability of pre-trained models. Our approach is grounded in viewing transferability as a measure of how easily a pre-trained model's representations can be trained to separate target classes, providing a unique perspective on transferability estimation. We rigorously evaluate the proposed metrics against state-of-the-art alternatives across diverse problem settings, demonstrating their robustness and practical utility. Additionally, we present theoretical insights that explain our metrics' efficacy and adaptability to various scenarios. We experimentally show that our metrics increase Kendall's Tau by up to 32% compared to the state-of-the-art baselines.
Estimation of Food Intake Quantity Using Inertial Signals from Smartwatches
Levi, Ioannis, Kyritsis, Konstantinos, Papapanagiotou, Vasileios, Tsakiridis, Georgios, Delopoulos, Anastasios
Accurate monitoring of eating behavior is crucial for managing obesity and eating disorders such as bulimia nervosa. At the same time, existing methods rely on multiple and/or specialized sensors, greatly harming adherence and ultimately, the quality and continuity of data. This paper introduces a novel approach for estimating the weight of a bite, from a commercial smartwatch. Our publicly-available dataset contains smartwatch inertial data from ten participants, with manually annotated start and end times of each bite along with their corresponding weights from a smart scale, under semi-controlled conditions. The proposed method combines extracted behavioral features such as the time required to load the utensil with food, with statistical features of inertial signals, that serve as input to a Support Vector Regression model to estimate bite weights. Under a leave-one-subject-out cross-validation scheme, our approach achieves a mean absolute error (MAE) of 3.99 grams per bite. To contextualize this performance, we introduce the improvement metric, that measures the relative MAE difference compared to a baseline model. Our method demonstrates a 17.41% improvement, while the adapted state-of-the art method shows a -28.89% performance against that same baseline. The results presented in this work establish the feasibility of extracting meaningful bite weight estimates from commercial smartwatch inertial sensors alone, laying the groundwork for future accessible, non-invasive dietary monitoring systems.
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
Choi, Kwanghee, Yeo, Eunjung, Chang, Kalvin, Watanabe, Shinji, Mortensen, David
Allophony refers to the variation in the phonetic realization of a phoneme based on its phonetic environment. Modeling allophones is crucial for atypical pronunciation assessment, which involves distinguishing atypical from typical pronunciations. However, recent phoneme classifier-based approaches often simplify this by treating various realizations as a single phoneme, bypassing the complexity of modeling allophonic variation. Motivated by the acoustic modeling capabilities of frozen self-supervised speech model (S3M) features, we propose MixGoP, a novel approach that leverages Gaussian mixture models to model phoneme distributions with multiple subclusters. Our experiments show that MixGoP achieves state-of-the-art performance across four out of five datasets, including dysarthric and non-native speech. Our analysis further suggests that S3M features capture allophonic variation more effectively than MFCCs and Mel spectrograms, highlighting the benefits of integrating MixGoP with S3M features.
Machine Learning for Everyone: Simplifying Healthcare Analytics with BigQuery ML
Salari, Mohammad Amir, Rahmani, Bahareh
The application of AI in healthcare allows for the identification of complex patterns in patient data, improving diagnostic accuracy, treatment personalization, and operational efficiency [1]. Healthcare providers are increasingly leveraging predictive analytics to foresee health outcomes, enabling earlier interventions and more targeted care [2][26]. For instance, AI models have proven effective in identifying high-risk patients and optimizing preventive care strategies [3]. Diabetes, a major global health challenge, requires early detection and preventive care. Predictive models built using accessible tools like BigQuery ML can help healthcare professionals identify at-risk individuals efficiently. Cloud computing serves as a critical tool for AI and ML in healthcare, addressing many of the technical and infrastructural challenges associated with large-scale data analysis. With scalable infrastructure, cloud platforms allow healthcare providers to process and store vast amounts of data, facilitating AI-driven insights without the need of extensive on-site resources [4].
DROP: Poison Dilution via Knowledge Distillation for Federated Learning
Syros, Georgios, Suri, Anshuman, Koushanfar, Farinaz, Nita-Rotaru, Cristina, Oprea, Alina
Federated Learning is vulnerable to adversarial manipulation, where malicious clients can inject poisoned updates to influence the global model's behavior. While existing defense mechanisms have made notable progress, they fail to protect against adversaries that aim to induce targeted backdoors under different learning and attack configurations. To address this limitation, we introduce DROP (Distillation-based Reduction Of Poisoning), a novel defense mechanism that combines clustering and activity-tracking techniques with extraction of benign behavior from clients via knowledge distillation to tackle stealthy adversaries that manipulate low data poisoning rates and diverse malicious client ratios within the federation. Through extensive experimentation, our approach demonstrates superior robustness compared to existing defenses across a wide range of learning configurations. Finally, we evaluate existing defenses and our method under the challenging setting of non-IID client data distribution and highlight the challenges of designing a resilient FL defense in this setting.
Advancing Precision Oncology Through Modeling of Longitudinal and Multimodal Data
Zhuang, Luoting, Park, Stephen H., Skates, Steven J., Prosper, Ashley E., Aberle, Denise R., Hsu, William
Cancer evolves continuously over time through a complex interplay of genetic, epigenetic, microenvironmental, and phenotypic changes. This dynamic behavior drives uncontrolled cell growth, metastasis, immune evasion, and therapy resistance, posing challenges for effective monitoring and treatment. However, today's data-driven research in oncology has primarily focused on cross-sectional analysis using data from a single modality, limiting the ability to fully characterize and interpret the disease's dynamic heterogeneity. Advances in multiscale data collection and computational methods now enable the discovery of longitudinal multimodal biomarkers for precision oncology. Longitudinal data reveal patterns of disease progression and treatment response that are not evident from single-timepoint data, enabling timely abnormality detection and dynamic treatment adaptation. Multimodal data integration offers complementary information from diverse sources for more precise risk assessment and targeting of cancer therapy. In this review, we survey methods of longitudinal and multimodal modeling, highlighting their synergy in providing multifaceted insights for personalized care tailored to the unique characteristics of a patient's cancer. We summarize the current challenges and future directions of longitudinal multimodal analysis in advancing precision oncology.
RSAttAE: An Information-Aware Attention-based Autoencoder Recommender System
Taromi, Amirhossein Dadashzadeh, Heydari, Sina, Hooshmand, Mohsen, Ramezani, Majid
Recommender systems play a crucial role in modern life, including information retrieval, the pharmaceutical industry, retail, and entertainment. The entertainment sector, in particular, attracts significant attention and generates substantial profits. This work proposes a new method for predicting unknown user-movie ratings to enhance customer satisfaction. To achieve this, we utilize the MovieLens 100K dataset. Our approach introduces an attention-based autoencoder to create meaningful representations and the XGBoost method for rating predictions. The results demonstrate that our proposal outperforms most of the existing state-of-the-art methods. Availability: github.com/ComputationIASBS/RecommSys
Multi-label Scandinavian Language Identification (SLIDE)
Fedorova, Mariia, Frydenberg, Jonas Sebulon, Handford, Victoria, Langø, Victoria Ovedie Chruickshank, Willoch, Solveig Helene, Midtgaard, Marthe Løken, Scherrer, Yves, Mæhlum, Petter, Samuel, David
Identifying closely related languages at sentence level is difficult, in particular because it is often impossible to assign a sentence to a single language. In this paper, we focus on multi-label sentence-level Scandinavian language identification (LID) for Danish, Norwegian Bokm\r{a}l, Norwegian Nynorsk, and Swedish. We present the Scandinavian Language Identification and Evaluation, SLIDE, a manually curated multi-label evaluation dataset and a suite of LID models with varying speed-accuracy tradeoffs. We demonstrate that the ability to identify multiple languages simultaneously is necessary for any accurate LID method, and present a novel approach to training such multi-label LID models.
EquiTabPFN: A Target-Permutation Equivariant Prior Fitted Networks
Arbel, Michael, Salinas, David, Hutter, Frank
However, these models overlook However, row-order symmetry is not the only symmetry a crucial equivariance property: the arbitrary relevant to tabular data. Another key symmetry pertains ordering of target dimensions should not influence to feature order, where the arrangement of columns should model predictions. In this study, we identify not influence model predictions. Recent work (Müller et al., this oversight as a source of incompressible 2024; Hollmann et al., 2025) has addressed this challenge by error, termed the equivariance gap, which introduces employing bi-attention mechanisms similar to those studied instability in predictions. To mitigate these in earlier work (Kossen et al., 2022). This approach alternates issues, we propose a novel model designed to preserve attention over rows and columns, making the models equivariance across output dimensions. Our equivariant to feature permutations and better suited for experimental results indicate that our proposed handling another inherent symmetry of tabular data.
Conditional Distribution Quantization in Machine Learning
Delattre, Blaise, Delattre, Sylvain, Vérine, Alexandre, Allauzen, Alexandre
Conditional expectation E(Y | X) often fails This limitation has important implications for downstream to capture the complexity of multimodal conditional tasks, particularly in uncertainty quantification for image distributions L(Y | X). To address this, restoration models used in safety-critical domains such as we propose using n-point conditional quantizations--functional autonomous driving and biological imaging. Many existing mappings of X that are learnable approaches rely on per-pixel estimates such as variance via gradient descent--to approximate L(Y | heatmaps (Kendall and Gal, 2017) or confidence intervals X). This approach adapts Competitive Learning (Angelopoulos et al., 2022) to visualize uncertainty. Vector Quantization (CLVQ), tailored for conditional While these methods provide valuable insights, they can distributions. It goes beyond single-valued struggle to represent structured uncertainty, overlooking predictions by providing multiple representative spatial correlations between neighboring pixels.