Statistical Learning
Stein Variational Newton Neural Network Ensembles
Flöge, Klemens, Moeed, Mohammed Abdul, Fortuin, Vincent
Deep neural network ensembles are powerful tools for uncertainty quantification, which have recently been re-interpreted from a Bayesian perspective. However, current methods inadequately leverage second-order information of the loss landscape, despite the recent availability of efficient Hessian approximations. We propose a novel approximate Bayesian inference method that modifies deep ensembles to incorporate Stein Variational Newton updates. Our approach uniquely integrates scalable modern Hessian approximations, achieving faster convergence and more accurate posterior distribution approximations. We validate the effectiveness of our method on diverse regression and classification tasks, demonstrating superior performance with a significantly reduced number of training epochs compared to existing ensemble-based methods, while enhancing uncertainty quantification and robustness against overfitting.
Expressivity of deterministic quantum computation with one qubit
Deterministic quantum computation with one qubit (DQC1) is of significant theoretical and practical interest due to its computational advantages in certain problems, despite its subuniversality with limited quantum resources. In this work, we introduce parameterized DQC1 as a quantum machine learning model. We demonstrate that the gradient of the measurement outcome of a DQC1 circuit with respect to its gate parameters can be computed directly using the DQC1 protocol. This allows for gradient-based optimization of DQC1 circuits, positioning DQC1 as the sole quantum protocol for both training and inference. We then analyze the expressivity of the parameterized DQC1 circuits, characterizing the set of learnable functions, and show that DQC1-based machine learning (ML) is as powerful as quantum neural networks based on universal computation. Our findings highlight the potential of DQC1 as a practical and versatile platform for ML, capable of rivaling more complex quantum computing models while utilizing simpler quantum resources.
A Bayesian explanation of machine learning models based on modes and functional ANOVA
Most methods in explainable AI (XAI) focus on providing reasons for the prediction of a given set of features. However, we solve an inverse explanation problem, i.e., given the deviation of a label, find the reasons of this deviation. We use a Bayesian framework to recover the ``true'' features, conditioned on the observed label value. We efficiently explain the deviation of a label value from the mode, by identifying and ranking the influential features using the ``distances'' in the ANOVA functional decomposition. We show that the new method is more human-intuitive and robust than methods based on mean values, e.g., SHapley Additive exPlanations (SHAP values). The extra costs of solving a Bayesian inverse problem are dimension-independent.
Data-Driven Hierarchical Open Set Recognition
Hannum, Andrew, Conway, Max, Lopez, Mario, Harrison, André
This paper presents a novel data-driven hierarchical approach to open set recognition (OSR) for robust perception in robotics and computer vision, utilizing constrained agglomerative clustering to automatically build a hierarchy of known classes in embedding space without requiring manual relational information. The method, demonstrated on the Animals with Attributes 2 (AwA2) dataset, achieves competitive results with an AUC ROC score of 0.82 and utility score of 0.85, while introducing two classification approaches (score-based and traversal-based) and a new Concentration Centrality (CC) metric for measuring hierarchical classification consistency. Although not surpassing existing models in accuracy, the approach provides valuable additional information about unknown classes through automatically generated hierarchies, requires no supplementary information beyond typical supervised model requirements, and introduces the Class Concentration Centrality (CCC) metric for evaluating unknown class placement consistency, with future work aimed at improving accuracy, validating the CC metric, and expanding to Large-Scale Open-Set Classification Protocols for ImageNet.
ELU-GCN: Effectively Label-Utilizing Graph Convolutional Network
Huang, Jincheng, Mo, Yujie, Shi, Xiaoshuang, Feng, Lei, Zhu, Xiaofeng
The message-passing mechanism of graph convolutional networks (i.e., GCNs) enables label information to be propagated to a broader range of neighbors, thereby increasing the utilization of labels. However, the label information is not always effectively utilized in the traditional GCN framework. To address this issue, we propose a new two-step framework called ELU-GCN. In the first stage, ELU-GCN conducts graph learning to learn a new graph structure (\ie ELU-graph), which enables GCNs to effectively utilize label information. In the second stage, we design a new graph contrastive learning on the GCN framework for representation learning by exploring the consistency and mutually exclusive information between the learned ELU graph and the original graph. Moreover, we theoretically demonstrate that the proposed method can ensure the generalization ability of GCNs. Extensive experiments validate the superiority of the proposed method.
Counterfactual Explanations via Riemannian Latent Space Traversal
Pegios, Paraskevas, Feragen, Aasa, Hansen, Andreas Abildtrup, Arvanitidis, Georgios
The adoption of increasingly complex deep models has fueled an urgent need for insight into how these models make predictions. Counterfactual explanations form a powerful tool for providing actionable explanations to practitioners. Previously, counterfactual explanation methods have been designed by traversing the latent space of generative models. Yet, these latent spaces are usually greatly simplified, with most of the data distribution complexity contained in the decoder rather than the latent embedding. Thus, traversing the latent space naively without taking the nonlinear decoder into account can lead to unnatural counterfactual trajectories. We introduce counterfactual explanations obtained using a Riemannian metric pulled back via the decoder and the classifier under scrutiny. This metric encodes information about the complex geometric structure of the data and the learned representation, enabling us to obtain robust counterfactual trajectories with high fidelity, as demonstrated by our experiments in real-world tabular datasets.
SIRA: Scalable Inter-frame Relation and Association for Radar Perception
Yataka, Ryoma, Wang, Pu Perry, Boufounos, Petros, Takahashi, Ryuhei
Conventional radar feature extraction faces limitations due to low spatial resolution, noise, multipath reflection, the presence of ghost targets, and motion blur. Such limitations can be exacerbated by nonlinear object motion, particularly from an ego-centric viewpoint. It becomes evident that to address these challenges, the key lies in exploiting temporal feature relation over an extended horizon and enforcing spatial motion consistency for effective association. To this end, this paper proposes SIRA (Scalable Inter-frame Relation and Association) with two designs. First, inspired by Swin Transformer, we introduce extended temporal relation, generalizing the existing temporal relation layer from two consecutive frames to multiple inter-frames with temporally regrouped window attention for scalability. Second, we propose motion consistency track with the concept of a pseudo-tracklet generated from observational data for better trajectory prediction and subsequent object association. Our approach achieves 58.11 mAP@0.5 for oriented object detection and 47.79 MOTA for multiple object tracking on the Radiate dataset, surpassing previous state-of-the-art by a margin of +4.11 mAP@0.5 and +9.94 MOTA, respectively.
Towards certification: A complete statistical validation pipeline for supervised learning in industry
Lacasa, Lucas, Pardo, Abel, Arbelo, Pablo, Sánchez, Miguel, Yeste, Pablo, Bascones, Noelia, Martínez-Cava, Alejandro, Rubio, Gonzalo, Gómez, Ignacio, Valero, Eusebio, de Vicente, Javier
The field of Machine Learning (ML) [1, 2] and its broad spectrum of applications has revolutionized a plethora of technological industries in recent years ranging from the energy sector or material sciences to telecommunications, finance or consumer goods, to cite some [3]. In the context of aeronautical engineering and aerospace technologies, the field has embraced ML tools only in recent years, and impact is growing at a rapid pace, ranging from generalpurpose ML-based fluid mechanics [4-6], aeroacoustics [7], wind turbines [8] or aerostructures [9] (including prediction of landing gear loads [10]) to flight trajectories optimization [11] or enhancing predictive maintenance [12, 13]: see the recent and illuminating reviews [14, 15] and references therein. Interestingly, the integration of ML-related tools and ideas in the aeronautical and aerospace industries is still in its infancy. Part of the reason is that any new technology has a necessary adoption curve [16, 17], and the fact that ML-solutions require expert knowledge at the crossroads of computer science and statistics -and a sophisticated operationalization infrastructure (MLOps) [18] - does not facilitate this adoption. However, a deeper reason is probably impeding faster adoption: while ML-technologies promise high performance and reduction in development and operating costs [19] (e.g. by reducing costs related to expensive and lengthy wind tunnel experiments and numerical simulations), ensuring adequate safety remains paramount in aeronautical industries, and ML-based tools are often seen as sophisticated black-boxes that suffer from low degree of trustability, and thus difficult to validate their safety. Therefore, air safety authorities demand rigorous validation and verification processes for these models, and industry leaders have started to propose guidelines and a roadmap on concepts of design assurance for neural network-related technologies [20-22]. However, only very recently industry has started to embrace the complexities of certifying ML models [23-27], prompting the initiation of discussions around the development of guidelines and a roadmap for design assurance, especially concerning network-related technologies. This pressing need underscores the imperative for collaborative efforts within the industry to establish robust validation frameworks that not only meet regulatory standards but also address the evolving challenges posed by ML integration. This has indeed been well understood and undertaken by Airbus who has established an internal working group on verification and validation of surrogate models in the frame of loads and stress domains.
R+R:Understanding Hyperparameter Effects in DP-SGD
Morsbach, Felix, Reubold, Jan, Strufe, Thorsten
Research on the effects of essential hyperparameters of DP-SGD lacks consensus, verification, and replication. Contradictory and anecdotal statements on their influence make matters worse. While DP-SGD is the standard optimization algorithm for privacy-preserving machine learning, its adoption is still commonly challenged by low performance compared to non-private learning approaches. As proper hyperparameter settings can improve the privacy-utility trade-off, understanding the influence of the hyperparameters promises to simplify their optimization towards better performance, and likely foster acceptance of private learning. To shed more light on these influences, we conduct a replication study: We synthesize extant research on hyperparameter influences of DP-SGD into conjectures, conduct a dedicated factorial study to independently identify hyperparameter effects, and assess which conjectures can be replicated across multiple datasets, model architectures, and differential privacy budgets. While we cannot (consistently) replicate conjectures about the main and interaction effects of the batch size and the number of epochs, we were able to replicate the conjectured relationship between the clipping threshold and learning rate. Furthermore, we were able to quantify the significant importance of their combination compared to the other hyperparameters.
UnSegMedGAT: Unsupervised Medical Image Segmentation using Graph Attention Networks Clustering
Adityaja, A. Mudit, Shigwan, Saurabh J., Kumar, Nitin
The data-intensive nature of supervised classification drives the interest of the researchers towards unsupervised approaches, especially for problems such as medical image segmentation, where labeled data is scarce. Building on the recent advancements of Vision transformers (ViT) in computer vision, we propose an unsupervised segmentation framework using a pre-trained Dino-ViT. In the proposed method, we leverage the inherent graph structure within the image to realize a significant performance gain for segmentation in medical images. For this, we introduce a modularity-based loss function coupled with a Graph Attention Network (GAT) to effectively capture the inherent graph topology within the image. Our method achieves state-of-the-art performance, even significantly surpassing or matching that of existing (semi)supervised technique such as MedSAM which is a Segment Anything Model in medical images. We demonstrate this using two challenging medical image datasets ISIC-2018 and CVC-ColonDB. This work underscores the potential of unsupervised approaches in advancing medical image analysis in scenarios where labeled data is scarce. The github repository of the code is available on [https://github.com/mudit-adityaja/UnSegMedGAT].