Goto

Collaborating Authors

 Statistical Learning


Review for NeurIPS paper: Federated Principal Component Analysis

Neural Information Processing Systems

Four knowledgeable reviewers support acceptance of this paper, in view of the novelty of the differentially private, federated PCA algorithm and the strength of the analysis provided for its variants. I concur with the reviewers. The reviewers pointed out issues with the clarity of the paper, and the authors promised several edits to address this issue in their rebuttal; please implement these changes.


Review for NeurIPS paper: Benchmarking Deep Learning Interpretability in Time Series Predictions

Neural Information Processing Systems

This work introduces a bunch of benchmarks for evaluating time series saliency methods (with respective metrics). The authors do a number of empirical evaluations, draw some conclusions about why certain things don't work, and propose a new saliency method based on that. There are a number of things that I like about this work and that was pointed out by the reviewers as well: there is a definite lack of datasets with groundtruth saliency in them so coming up with such a dataset (and associated metrics) is a worthy contribution by itself (though perhaps not rising up to the bar of acceptance at NeurIPS). In general, everyone agreed that this part of the paper is good. What was more controversial: is the subsequent analysis interesting and novel enough?


Facies Classification with Copula Entropy

arXiv.org Artificial Intelligence

Facies are the type of rocks with similar characteristics given by geologists and facies classification is of very significance in geological tasks, such as formation evaluation, reservoir characterization. As the geological data accumulates, there are growing interests in facies classification with machine learning methods [1, 2, 3, 4, 5, 6, 7, 8, 9]. There are two issues with the existing works on facies classification. First, the machine learning models are built without variable selection or with only very primary method, such as cross-validation, which makes the classifiers with useless variable as inputs and therefore with low performance. Second, most of the models for facies classification are block-box, such as deep learning [5, 10, 11], Boostings or SVMs[7], which are un-interpretable to geologists. Variable selection is a common task that selects a subset from all the available variables for machine learning models. By this, the accuracy of the predictive models built with the selected variables can be improved compared with those built without selection. The traditional method for variable selection are mainly based on likelihoods, such as AIC, BIC, or accuracy, such as LASSO [12], or correlation, such as HSIC [13], distance correlation [14], and copula entropy [15]. Copula Entropy (CE) is a recently proposed rigorous mathematical concept for measuring multivariate statistical independence and is proved to be equivalent to mutual information in information theory [16].


Predictive Modeling and Uncertainty Quantification of Fatigue Life in Metal Alloys using Machine Learning

arXiv.org Artificial Intelligence

Recent advancements in machine learning-based methods have demonstrated great potential for improved property prediction in material science. However, reliable estimation of the confidence intervals for the predicted values remains a challenge, due to the inherent complexities in material modeling. This study introduces a novel approach for uncertainty quantification in fatigue life prediction of metal materials based on integrating knowledge from physics-based fatigue life models and machine learning models. The proposed approach employs physics-based input features estimated using the Basquin fatigue model to augment the experimentally collected data of fatigue life. Furthermore, a physics-informed loss function that enforces boundary constraints for the estimated fatigue life of considered materials is introduced for the neural network models. Experimental validation on datasets comprising collected data from fatigue life tests for Titanium alloys and Carbon steel alloys demonstrates the effectiveness of the proposed approach. The synergy between physics-based models and data-driven models enhances the consistency in predicted values and improves uncertainty interval estimates.


Reduced-order modeling and classification of hydrodynamic pattern formation in gravure printing

arXiv.org Artificial Intelligence

Hydrodynamic pattern formation phenomena in printing and coating processes are still not fully understood. However, fundamental understanding is essential to achieve high-quality printed products and to tune printed patterns according to the needs of a specific application like printed electronics, graphical printing, or biomedical printing. The aim of the paper is to develop an automated pattern classification algorithm based on methods from supervised machine learning and reduced-order modeling. We use the HYPA-p dataset, a large image dataset of gravure-printed images, which shows various types of hydrodynamic pattern formation phenomena. It enables the correlation of printing process parameters and resulting printed patterns for the first time. 26880 images of the HYPA-p dataset have been labeled by a human observer as dot patterns, mixed patterns, or finger patterns; 864000 images (97%) are unlabeled. A singular value decomposition (SVD) is used to find the modes of the labeled images and to reduce the dimensionality of the full dataset by truncation and projection. Selected machine learning classification techniques are trained on the reduced-order data. We investigate the effect of several factors, including classifier choice, whether or not fast Fourier transform (FFT) is used to preprocess the labeled images, data balancing, and data normalization. The best performing model is a k-nearest neighbor (kNN) classifier trained on unbalanced, FFT-transformed data with a test error of 3%, which outperforms a human observer by 7%. Data balancing slightly increases the test error of the kNN-model to 5%, but also increases the recall of the mixed class from 90% to 94%. Finally, we demonstrate how the trained models can be used to predict the pattern class of unlabeled images and how the predictions can be correlated to the printing process parameters, in the form of regime maps.


Decision-Focused Learning for Complex System Identification: HVAC Management System Application

arXiv.org Artificial Intelligence

As opposed to conventional training methods tailored to minimize a given statistical metric or task-agnostic loss (e.g., mean squared error), Decision-Focused Learning (DFL) trains machine learning models for optimal performance in downstream decision-making tools. We argue that DFL can be leveraged to learn the parameters of system dynamics, expressed as constraint of the convex optimization control policy, while the system control signal is being optimized, thus creating an end-to-end learning framework. This is particularly relevant for systems in which behavior changes once the control policy is applied, hence rendering historical data less applicable. The proposed approach can perform system identification - i.e., determine appropriate parameters for the system analytical model - and control simultaneously to ensure that the model's accuracy is focused on areas most relevant to control. Furthermore, because black-box systems are non-differentiable, we design a loss function that requires solely to measure the system response. We propose pre-training on historical data and constraint relaxation to stabilize the DFL and deal with potential infeasibilities in learning. We demonstrate the usefulness of the method on a building Heating, Ventilation, and Air Conditioning day-ahead management system for a realistic 15-zone building located in Denver, US. The results show that the conventional RC building model, with the parameters obtained from historical data using supervised learning, underestimates HVAC electrical power consumption. For our case study, the ex-post cost is on average six times higher than the expected one. Meanwhile, the same RC model with parameters obtained via DFL underestimates the ex-post cost only by 3%.


Whisper D-SGD: Correlated Noise Across Agents for Differentially Private Decentralized Learning

arXiv.org Artificial Intelligence

Decentralized learning enables distributed agents to train a shared machine learning model through local computation and peer-to-peer communication. Although each agent retains its dataset locally, the communication of local models can still expose private information to adversaries. To mitigate these threats, local differential privacy (LDP) injects independent noise per agent, but it suffers a larger utility gap than central differential privacy (CDP). We introduce Whisper D-SGD, a novel covariance-based approach that generates correlated privacy noise across agents, unifying several state-of-the-art methods as special cases. By leveraging network topology and mixing weights, Whisper D-SGD optimizes the noise covariance to achieve network-wide noise cancellation. Experimental results show that Whisper D-SGD cancels more noise than existing pairwise-correlation schemes, substantially narrowing the CDP-LDP gap and improving model performance under the same privacy guarantees.


Principal Graph Encoder Embedding and Principal Community Detection

arXiv.org Machine Learning

In this paper, we introduce the concept of principal communities and propose a principal graph encoder embedding method that concurrently detects these communities and achieves vertex embedding. Given a graph adjacency matrix with vertex labels, the method computes a sample community score for each community, ranking them to measure community importance and estimate a set of principal communities. The method then produces a vertex embedding by retaining only the dimensions corresponding to these principal communities. Theoretically, we define the population version of the encoder embedding and the community score based on a random Bernoulli graph distribution. We prove that the population principal graph encoder embedding preserves the conditional density of the vertex labels and that the population community score successfully distinguishes the principal communities. We conduct a variety of simulations to demonstrate the finite-sample accuracy in detecting ground-truth principal communities, as well as the advantages in embedding visualization and subsequent vertex classification. The method is further applied to a set of real-world graphs, showcasing its numerical advantages, including robustness to label noise and computational scalability.


Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval

arXiv.org Artificial Intelligence

Video surveillance systems are crucial components for ensuring public safety and management in smart city. As a fundamental task in video surveillance, text-to-image person retrieval aims to retrieve the target person from an image gallery that best matches the given text description. Most existing text-to-image person retrieval methods are trained in a supervised manner that requires sufficient labeled data in the target domain. However, it is common in practice that only unlabeled data is available in the target domain due to the difficulty and cost of data annotation, which limits the generalization of existing methods in practical application scenarios. To address this issue, we propose a novel unsupervised domain adaptation method, termed Graph-Based Cross-Domain Knowledge Distillation (GCKD), to learn the cross-modal feature representation for text-to-image person retrieval in a cross-dataset scenario. The proposed GCKD method consists of two main components. Firstly, a graph-based multi-modal propagation module is designed to bridge the cross-domain correlation among the visual and textual samples. Secondly, a contrastive momentum knowledge distillation module is proposed to learn the cross-modal feature representation using the online knowledge distillation strategy. By jointly optimizing the two modules, the proposed method is able to achieve efficient performance for cross-dataset text-to-image person retrieval. acExtensive experiments on three publicly available text-to-image person retrieval datasets demonstrate the effectiveness of the proposed GCKD method, which consistently outperforms the state-of-the-art baselines.


Statistical Verification of Linear Classifiers

arXiv.org Machine Learning

We propose a homogeneity test closely related to the concept of linear separability between two samples. Using the test one can answer the question whether a linear classifier is merely ``random'' or effectively captures differences between two classes. We focus on establishing upper bounds for the test's \emph{p}-value when applied to two-dimensional samples. Specifically, for normally distributed samples we experimentally demonstrate that the upper bound is highly accurate. Using this bound, we evaluate classifiers designed to detect ER-positive breast cancer recurrence based on gene pair expression. Our findings confirm significance of IGFBP6 and ELOVL5 genes in this process.