Statistical Learning
CEREAL: Few-Sample Clustering Evaluation
Nayak, Nihal V., Elenberg, Ethan R., Rosenbaum, Clemens
Evaluating clustering quality with reliable evaluation metrics like normalized mutual information (NMI) requires labeled data that can be expensive to annotate. We focus on the underexplored problem of estimating clustering quality with limited labels. We adapt existing approaches from the few-sample model evaluation literature to actively sub-sample, with a learned surrogate model, the most informative data points for annotation to estimate the evaluation metric. However, we find that their estimation can be biased and only relies on the labeled data. To that end, we introduce CEREAL, a comprehensive framework for few-sample clustering evaluation that extends active sampling approaches in three key ways. First, we propose novel NMI-based acquisition functions that account for the distinctive properties of clustering and uncertainties from a learned surrogate model. Next, we use ideas from semi-supervised learning and train the surrogate model with both the labeled and unlabeled data. Finally, we pseudo-label the unlabeled data with the surrogate model. We run experiments to estimate NMI in an active sampling pipeline on three datasets across vision and language. Our results show that CEREAL reduces the area under the absolute error curve by up to 57% compared to the best sampling baseline. We perform an extensive ablation study to show that our framework is agnostic to the choice of clustering algorithm and evaluation metric. We also extend CEREAL from clusterwise annotations to pairwise annotations. Overall, CEREAL can efficiently evaluate clustering with limited human annotations.
Matryoshka Representation Learning
Kusupati, Aditya, Bhatt, Gantavya, Rege, Aniket, Wallingford, Matthew, Sinha, Aditya, Ramanujan, Vivek, Howard-Snyder, William, Chen, Kaifeng, Kakade, Sham, Jain, Prateek, Farhadi, Ali
Learned representations are a central component in modern ML systems, serving a multitude of downstream tasks. When training such representations, it is often the case that computational and statistical constraints for each downstream task are unknown. In this context rigid, fixed capacity representations can be either over or under-accommodating to the task at hand. This leads us to ask: can we design a flexible representation that can adapt to multiple downstream tasks with varying computational resources? Our main contribution is Matryoshka Representation Learning (MRL) which encodes information at different granularities and allows a single embedding to adapt to the computational constraints of downstream tasks. MRL minimally modifies existing representation learning pipelines and imposes no additional cost during inference and deployment. MRL learns coarse-to-fine representations that are at least as accurate and rich as independently trained low-dimensional representations. The flexibility within the learned Matryoshka Representations offer: (a) up to 14x smaller embedding size for ImageNet-1K classification at the same level of accuracy; (b) up to 14x real-world speed-ups for large-scale retrieval on ImageNet-1K and 4K; and (c) up to 2% accuracy improvements for long-tail few-shot classification, all while being as robust as the original representations. Finally, we show that MRL extends seamlessly to web-scale datasets (ImageNet, JFT) across various modalities -- vision (ViT, ResNet), vision + language (ALIGN) and language (BERT). MRL code and pretrained models are open-sourced at https://github.com/RAIVNLab/MRL.
Fault Prognosis in Particle Accelerator Power Electronics Using Ensemble Learning
Radaideh, Majdi I., Pappas, Chris, Wezensky, Mark, Ramuhalli, Pradeep, Cousineau, Sarah
Early fault detection and fault prognosis are crucial to ensure efficient and safe operations of complex engineering systems such as the Spallation Neutron Source (SNS) and its power electronics (high voltage converter modulators). Following an advanced experimental facility setup that mimics SNS operating conditions, the authors successfully conducted 21 fault prognosis experiments, where fault precursors are introduced in the system to a degree enough to cause degradation in the waveform signals, but not enough to reach a real fault. Nine different machine learning techniques based on ensemble trees, convolutional neural networks, support vector machines, and hierarchical voting ensembles are proposed to detect the fault precursors. Although all 9 models have shown a perfect and identical performance during the training and testing phase, the performance of most models has decreased in the prognosis phase once they got exposed to real-world data from the 21 experiments. The hierarchical voting ensemble, which features multiple layers of diverse models, maintains a distinguished performance in early detection of the fault precursors with 95% success rate (20/21 tests), followed by adaboost and extremely randomized trees with 52% and 48% success rates, respectively. The support vector machine models were the worst with only 24% success rate (5/21 tests). The study concluded that a successful implementation of machine learning in the SNS or particle accelerator power systems would require a major upgrade in the controller and the data acquisition system to facilitate streaming and handling big data for the machine learning models. In addition, this study shows that the best performing models were diverse and based on the ensemble concept to reduce the bias and hyperparameter sensitivity of individual models.
Identifying Latent Causal Content for Multi-Source Domain Adaptation
Liu, Yuhang, Zhang, Zhen, Gong, Dong, Gong, Mingming, Huang, Biwei, Zhang, Kun, Shi, Javen Qinfeng
Multi-source domain adaptation (MSDA) learns to predict the labels in target domain data, under the setting that data from multiple source domains are labelled and data from the target domain are unlabelled. Most methods for this task focus on learning invariant representations across domains. However, their success relies heavily on the assumption that the label distribution remains consistent across domains, which may not hold in general real-world problems. In this paper, we propose a new and more flexible assumption, termed \textit{latent covariate shift}, where a latent content variable $\mathbf{z}_c$ and a latent style variable $\mathbf{z}_s$ are introduced in the generative process, with the marginal distribution of $\mathbf{z}_c$ changing across domains and the conditional distribution of the label given $\mathbf{z}_c$ remaining invariant across domains. We show that although (completely) identifying the proposed latent causal model is challenging, the latent content variable can be identified up to scaling by using its dependence with labels from source domains, together with the identifiability conditions of nonlinear ICA. This motivates us to propose a novel method for MSDA, which learns the invariant label distribution conditional on the latent content variable, instead of learning invariant representations. Empirical evaluation on simulation and real data demonstrates the effectiveness of the proposed method.
Efficiently Learning Small Policies for Locomotion and Manipulation
Hegde, Shashank, Sukhatme, Gaurav S.
Neural control of memory-constrained, agile robots requires small, yet highly performant models. We leverage graph hyper networks to learn graph hyper policies trained with off-policy reinforcement learning resulting in networks that are two orders of magnitude smaller than commonly used networks yet encode policies comparable to those encoded by much larger networks trained on the same task. We show that our method can be appended to any off-policy reinforcement learning algorithm, without any change in hyperparameters, by showing results across locomotion and manipulation tasks. Further, we obtain an array of working policies, with differing numbers of parameters, allowing us to pick an optimal network for the memory constraints of a system. Training multiple policies with our method is as sample efficient as training a single policy. Finally, we provide a method to select the best architecture, given a constraint on the number of parameters. Project website: https://sites.google.com/usc.edu/graphhyperpolicy
Amplitude Scintillation Forecasting Using Bagged Trees
Darya, Abdollah Masoud, Al-Owais, Aisha Abdulla, Shaikh, Muhammad Mubasshir, Fernini, Ilias
Electron density irregularities present within the ionosphere induce significant fluctuations in global navigation satellite system (GNSS) signals. Fluctuations in signal power are referred to as amplitude scintillation and can be monitored through the S4 index. Forecasting the severity of amplitude scintillation based on historical S4 index data is beneficial when real-time data is unavailable. In this work, we study the possibility of using historical data from a single GPS scintillation monitoring receiver to train a machine learning (ML) model to forecast the severity of amplitude scintillation, either weak, moderate, or severe, with respect to temporal and spatial parameters. Six different ML models were evaluated and the bagged trees model was the most accurate among them, achieving a forecasting accuracy of $81\%$ using a balanced dataset, and $97\%$ using an imbalanced dataset.
Using Knowledge Distillation to improve interpretable models in a retail banking context
Biehler, Maxime, Guermazi, Mohamed, Starck, Célim
Although the banking sector holds massive troves of data regarding its customers, products and transactions, and is no stranger to using quantitative tools to inform its decisions, two constraints usually weigh on the development of predictive models. The first one lies in the regulatory obligation to use interpretable models for a wide range of issues, with the management function being able to explain both the way a model was trained and why specific decisions have been made. Indeed, the European Banking Authority (2020) urges banking institutions to "understand the models used, and their methodology, input data, assumptions, limitations and outputs". The second has to do with the production environments available to deploy the models on. Due to the persistence of legacy systems, cost constraints or execution time limits -- think real time e-commerce fraud detection -- models may be limited to simple operations and conditions, i.e. a set of rules rather than a random forest, light computations in place of a fully fledged neural network. Modeling for retail banking use cases means dealing with both these strong customers protections -- enforced through regular audits -- and the high data volume which at times shortens the time allocated to each sample. These shackles help explain why modeling practices in retail banking departments are centered around simple and interpretable models such as the logistic regression or the (shallow) decision tree.
Deep Recurrent Encoder: A scalable end-to-end network to model brain signals
Chehab, Omar, Defossez, Alexandre, Loiseau, Jean-Christophe, Gramfort, Alexandre, King, Jean-Remi
A major goal of cognitive neuroscience consists of identifying how the brain responds to distinct experimental conditions. While descriptive statistics and statistical tests are classically used to analyze neural data [1], this approach is not suited to predict how the brain should react to new conditions. The resulting models of the brain can thus be particularly challenging to compare. By contrast, predictive encoding models [2, 3] can be directly trained to predict brain responses to various experimental conditions, and compared on their ability to accurately predict novel conditions. For example, encoding models allow the estimation of integration constants in the brain [4, 5], the hierarchical organization of visual [6] and speech processing [7, 8]. Beyond MEG, predictive models have enabled automatic segmentation [9] and dynamical system identification [10, 11]. In functional Magnetic Resonance Imaging, predictive encoding models are starting to emulate complex neural processing [12] and are a step towards discovering new phenomena [13, 14]. Yet, this general objective of developing encoding models faces three major challenges when working with non-invasive and time-resolved signals collected by magneto-and electro-encephalography (M/EEG).
Paralinguistic Privacy Protection at the Edge
Aloufi, Ranya, Haddadi, Hamed, Boyle, David
Voice user interfaces and digital assistants are rapidly entering our lives and becoming singular touch points spanning our devices. These always-on services capture and transmit our audio data to powerful cloud services for further processing and subsequent actions. Our voices and raw audio signals collected through these devices contain a host of sensitive paralinguistic information that is transmitted to service providers regardless of deliberate or false triggers. As our emotional patterns and sensitive attributes like our identity, gender, well-being, are easily inferred using deep acoustic models, we encounter a new generation of privacy risks by using these services. One approach to mitigate the risk of paralinguistic-based privacy breaches is to exploit a combination of cloud-based processing with privacy-preserving, on-device paralinguistic information learning and filtering before transmitting voice data. In this paper we introduce EDGY, a configurable, lightweight, disentangled representation learning framework that transforms and filters high-dimensional voice data to identify and contain sensitive attributes at the edge prior to offloading to the cloud. We evaluate EDGY's on-device performance and explore optimization techniques, including model quantization and knowledge distillation, to enable private, accurate and efficient representation learning on resource-constrained devices. Our results show that EDGY runs in tens of milliseconds with 0.2% relative improvement in "zero-shot" ABX score or minimal performance penalties of approximately 5.95% word error rate (WER) in learning linguistic representations from raw voice signals, using a CPU and a single-core ARM processor without specialized hardware.
PL-kNN: A Parameterless Nearest Neighbors Classifier
Jodas, Danilo Samuel, Passos, Leandro Aparecido, Adeel, Ahsan, Papa, João Paulo
Demands for minimum parameter setup in machine learning models are desirable to avoid time-consuming optimization processes. The $k$-Nearest Neighbors is one of the most effective and straightforward models employed in numerous problems. Despite its well-known performance, it requires the value of $k$ for specific data distribution, thus demanding expensive computational efforts. This paper proposes a $k$-Nearest Neighbors classifier that bypasses the need to define the value of $k$. The model computes the $k$ value adaptively considering the data distribution of the training set. We compared the proposed model against the standard $k$-Nearest Neighbors classifier and two parameterless versions from the literature. Experiments over 11 public datasets confirm the robustness of the proposed approach, for the obtained results were similar or even better than its counterpart versions.