Statistical Learning
Computer-aided shape features extraction and regression models for predicting the ascending aortic aneurysm growth rate
Geronzi, Leonardo, Martinez, Antonio, Rochette, Michel, Yan, Kexin, Bel-Brunon, Aline, Haigron, Pascal, Escrig, Pierre, Tomasi, Jacques, Daniel, Morgan, Lalande, Alain, Lin, Siyu, Marin-Castrillon, Diana Marcela, Bouchot, Olivier, Porterie, Jean, Valentini, Pier Paolo, Biancolini, Marco Evangelos
Objective: ascending aortic aneurysm growth prediction is still challenging in clinics. In this study, we evaluate and compare the ability of local and global shape features to predict ascending aortic aneurysm growth. Material and methods: 70 patients with aneurysm, for which two 3D acquisitions were available, are included. Following segmentation, three local shape features are computed: (1) the ratio between maximum diameter and length of the ascending aorta centerline, (2) the ratio between the length of external and internal lines on the ascending aorta and (3) the tortuosity of the ascending tract. By exploiting longitudinal data, the aneurysm growth rate is derived. Using radial basis function mesh morphing, iso-topological surface meshes are created. Statistical shape analysis is performed through unsupervised principal component analysis (PCA) and supervised partial least squares (PLS). Two types of global shape features are identified: three PCA-derived and three PLS-based shape modes. Three regression models are set for growth prediction: two based on gaussian support vector machine using local and PCA-derived global shape features; the third is a PLS linear regression model based on the related global shape features. The prediction results are assessed and the aortic shapes most prone to growth are identified. Results: the prediction root mean square error from leave-one-out cross-validation is: 0.112 mm/month, 0.083 mm/month and 0.066 mm/month for local, PCA-based and PLS-derived shape features, respectively. Aneurysms close to the root with a large initial diameter report faster growth. Conclusion: global shape features might provide an important contribution for predicting the aneurysm growth.
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
Viel, Stefano, Viano, Luca, Cevher, Volkan
This paper introduces the SOAR framework for imitation learning. SOAR is an algorithmic template that learns a policy from expert demonstrations with a primal dual style algorithm that alternates cost and policy updates. Within the policy updates, the SOAR framework uses an actor critic method with multiple critics to estimate the critic uncertainty and build an optimistic critic fundamental to drive exploration. When instantiated in the tabular setting, we get a provable algorithm with guarantees that matches the best known results in $\epsilon$. Practically, the SOAR template is shown to boost consistently the performance of imitation learning algorithms based on Soft Actor Critic such as f-IRL, ML-IRL and CSIL in several MuJoCo environments. Overall, thanks to SOAR, the required number of episodes to achieve the same performance is reduced by half.
Meta-Learning to Explore via Memory Density Feedback
Exploration algorithms for reinforcement learning typically replace or augment the reward function with an additional "intrinsic" reward that trains the agent to seek previously unseen states of the environment. Here, we consider an exploration algorithm that exploits meta-learning, or learning to learn, such that the agent learns to maximize its exploration progress within a single episode, even between epochs of training. The agent learns a policy that aims to minimize the probability density of new observations with respect to all of its memories. In addition, it receives as feedback evaluations of the current observation density and retains that feedback in a recurrent network. By remembering trajectories of density, the agent learns to navigate a complex and growing landscape of familiarity in real-time, allowing it to maximize its exploration progress even in completely novel states of the environment for which its policy has not been trained. Introduction In reinforcement learning (RL), exploration refers to algorithms that induce an agent to observe as much of a given task as possible. All RL algorithms include some form of random exploration, such as the epsilon-greedy policy or by additionally training to maximize the policy's entropy. These algorithms are necessary for the agent to find rewarding states and expand its policy, but often fall short when rewards are sparsely distributed, that is, requiring non-obvious and improbable sequences of action.
A Minimalist Example of Edge-of-Stability and Progressive Sharpening
Liu, Liming, Zhang, Zixuan, Du, Simon, Zhao, Tuo
Recent advances in deep learning optimization have unveiled two intriguing phenomena under large learning rates: Edge of Stability (EoS) and Progressive Sharpening (PS), challenging classical Gradient Descent (GD) analyses. Current research approaches, using either generalist frameworks or minimalist examples, face significant limitations in explaining these phenomena. This paper advances the minimalist approach by introducing a two-layer network with a two-dimensional input, where one dimension is relevant to the response and the other is irrelevant. Through this model, we rigorously prove the existence of progressive sharpening and self-stabilization under large learning rates, and establish non-asymptotic analysis of the training dynamics and sharpness along the entire GD trajectory. Besides, we connect our minimalist example to existing works by reconciling the existence of a well-behaved ``stable set" between minimalist and generalist analyses, and extending the analysis of Gradient Flow Solution sharpness to our two-dimensional input scenario. These findings provide new insights into the EoS phenomenon from both parameter and input data distribution perspectives, potentially informing more effective optimization strategies in deep learning practice.
Inductive randomness predictors
This paper introduces inductive randomness predictors, which form a superset of inductive conformal predictors. Its focus is on a very simple special case, binary inductive randomness predictors. It is interesting that binary inductive randomness predictors have an advantage over inductive conformal predictors, although they also have a serious disadvantage. This advantage will allow us to reach the surprising conclusion that non-trivial inductive conformal predictors are inadmissible in the sense of statistical decision theory.
Seeded Poisson Factorization: Leveraging domain knowledge to fit topic models
Prostmaier, Bernd, Vรกvra, Jan, Grรผn, Bettina, Hofmarcher, Paul
Topic models are widely used for discovering latent thematic structures in large text corpora, yet traditional unsupervised methods often struggle to align with predefined conceptual domains. This paper introduces Seeded Poisson Factorization (SPF), a novel approach that extends the Poisson Factorization framework by incorporating domain knowledge through seed words. SPF enables a more interpretable and structured topic discovery by modifying the prior distribution of topic-specific term intensities, assigning higher initial rates to predefined seed words. The model is estimated using variational inference with stochastic gradient optimization, ensuring scalability to large datasets. We apply SPF to an Amazon customer feedback dataset, leveraging predefined product categories as guiding structures. Our evaluation demonstrates that SPF achieves superior classification performance compared to alternative guided topic models, particularly in terms of computational efficiency and predictive performance. Furthermore, robustness checks highlight SPF's ability to adaptively balance domain knowledge and data-driven topic discovery, even in cases of imperfect seed word selection. These results establish SPF as a powerful and scalable alternative for integrating expert knowledge into topic modeling, enhancing both interpretability and efficiency in real-world applications.
Clustered KL-barycenter design for policy evaluation
Weissmann, Simon, Freihaut, Till, Vernade, Claire, Ramponi, Giorgia, Dรถring, Leif
In the context of stochastic bandit models, this article examines how to design sample-efficient behavior policies for the importance sampling evaluation of multiple target policies. From importance sampling theory, it is well established that sample efficiency is highly sensitive to the KL divergence between the target and importance sampling distributions. We first analyze a single behavior policy defined as the KL-barycenter of the target policies. Then, we refine this approach by clustering the target policies into groups with small KL divergences and assigning each cluster its own KL-barycenter as a behavior policy. This clustered KL-based policy evaluation (CKL-PE) algorithm provides a novel perspective on optimal policy selection. We prove upper bounds on the sample complexity of our method and demonstrate its effectiveness with numerical validation.
Joint Tensor and Inter-View Low-Rank Recovery for Incomplete Multiview Clustering
Wang, Jianyu, Zhao, Zhengqiao, Dobigeon, Nicolas, Chen, Jingdong
ULTIVIEW data consists of samples captured from multiple perspectives or modalities [1], making it wellsuited the application of MVC when some samples are missing in one for classification and clustering analysis. It has important or more views. In fact, in real-world applications, it is often applications in fields such as image analysis [2], video difficult to obtain the complete data for all views of interest due face recognition [3], and bioinformatics [4]. Compared to to data collection limitations such as sensor failures, data corruption, single-view approaches, which represent only one perspective or interrupted data acquisition processes. As a result, and often provide a limited understanding of objects, incomplete multiview clustering (IMVC) algorithms are drawing multiview clustering (MVC) methods leverage complementary increasing attention [9], [10]. In IMVC, representations information from different views to obtain a more comprehensive from different views are often partially available, resulting in and robust representation of the data [5]. By imposing key the loss of crucial information and difficulty in aligning views, assumptions such as independence or correlation among different significantly impacting clustering performance. The main challenge views, MVC has shown to enhance clustering performance of IMVC lies in effectively utilizing the available data by modeling deeper structures across views, overcoming the across all views while handling the missing samples [11], [12].
AILS-NTUA at SemEval-2025 Task 4: Parameter-Efficient Unlearning for Large Language Models using Data Chunking
Premptis, Iraklis, Lymperaiou, Maria, Filandrianos, Giorgos, Mastromichalakis, Orfeas Menis, Voulodimos, Athanasios, Stamou, Giorgos
The Unlearning Sensitive Content from Large Language Models task aims to remove targeted datapoints from trained models while minimally affecting their general knowledge. In our work, we leverage parameter-efficient, gradient-based unlearning using low-rank (LoRA) adaptation and layer-focused fine-tuning. To further enhance unlearning effectiveness, we employ data chunking, splitting forget data into disjoint partitions and merging them with cyclically sampled retain samples at a pre-defined ratio. Our task-agnostic method achieves an outstanding forget-retain balance, ranking first on leaderboards and significantly outperforming baselines and competing systems.
A Binary Classification Social Network Dataset for Graph Machine Learning
Ali, Adnan, Li, Jinglong, Chen, Huanhuan, Ajlouni, AlMotasem Bellah Al
Social networks have a vast range of applications with graphs. The available benchmark datasets are citation, co-occurrence, e-commerce networks, etc, with classes ranging from 3 to 15. However, there is no benchmark classification social network dataset for graph machine learning. This paper fills the gap and presents the Binary Classification Social Network Dataset (\textit{BiSND}), designed for graph machine learning applications to predict binary classes. We present the BiSND in \textit{tabular and graph} formats to verify its robustness across classical and advanced machine learning. We employ a diverse set of classifiers, including four traditional machine learning algorithms (Decision Trees, K-Nearest Neighbour, Random Forest, XGBoost), one Deep Neural Network (multi-layer perceptrons), one Graph Neural Network (Graph Convolutional Network), and three state-of-the-art Graph Contrastive Learning methods (BGRL, GRACE, DAENS). Our findings reveal that BiSND is suitable for classification tasks, with F1-scores ranging from 67.66 to 70.15, indicating promising avenues for future enhancements.