Learning Graphical Models
Learning User Embeddings from Temporal Social Media Data: A Survey
Hasan, Fatema, Xu, Kevin S., Foulds, James R., Pan, Shimei
User-generated data on social media contain rich information about who we are, what we like and how we make decisions. In this paper, we survey representative work on learning a concise latent user representation (a.k.a. user embedding) that can capture the main characteristics of a social media user. The learned user embeddings can later be used to support different downstream user analysis tasks such as personality modeling, suicidal risk assessment and purchase decision prediction. The temporal nature of user-generated data on social media has largely been overlooked in much of the existing user embedding literature. In this survey, we focus on research that bridges the gap by incorporating temporal/sequential information in user representation learning. We categorize relevant papers along several key dimensions, identify limitations in the current work and suggest future research directions.
Deep Multistage Multi-Task Learning for Quality Prediction of Multistage Manufacturing Systems
Yan, Hao, Sergin, Nurretin Dorukhan, Brenneman, William A., Lange, Stephen Joseph, Ba, Shan
In multistage manufacturing systems, modeling multiple quality indices based on the process sensing variables is important. However, the classic modeling technique predicts each quality variable one at a time, which fails to consider the correlation within or between stages. We propose a deep multistage multi-task learning framework to jointly predict all output sensing variables in a unified end-to-end learning framework according to the sequential system architecture in the MMS. Our numerical studies and real case study have shown that the new model has a superior performance compared to many benchmark methods as well as great interpretability through developed variable selection techniques.
Posterior Regularisation on Bayesian Hierarchical Mixture Clustering
Huang, Weipeng, Ng, Tin Lok James, Laitonjam, Nishma, Hurley, Neil J.
The framework is founded on an approach of minimising the Kullback-Leibler (KL) divergence between a variational solution and the posterior, in a constrained space. The works (Dudík et al., 2004, 2007; Altun and Smola, 2006) first raised the idea of including constraints in maximum entropy density estimation and provided a theoretical analysis. Based on convex duality theory, the optimal solution of the regularised posterior is found to be the original posterior of the model, discounted by the constrained pseudo likelihood introduced by the constraints. Later work founded on the idea of posterior constraints includes (Graça et al., 2009) which proposed constraining the E-step of an Expectation-maximization (EM) algorithm, in order to impose feature constraints on the solution.
Amazon.com: Probability and Statistics for Data Science: Math + R + Data (Chapman & Hall/CRC Data Science Series) (9781138393295): Matloff, Norman: Books
I believe that the book describes itself quite well when it says: Mathematically correct yet highly intuitive…This book would be great for a class that one takes before one takes my statistical learning class. I often run into beginning graduate Data Science students whose background is not math (e.g., CS or Business) and they are not ready…The book fills an important niche, in that it provides a self-contained introduction to material that is useful for a higher-level statistical learning course. I think that it compares well with competing books, particularly in that it takes a more "Data Science" and "example driven" approach than more classical books." "This text by Matloff (Univ. of California, Davis) affords an excellent introduction to statistics for the data science student…Its examples are often drawn from data science applications such as hidden Markov models and remote sensing, to name a few… All the models and concepts are explained well in precise mathematical terms (not presented as formal proofs), to help students gain an intuitive understanding."
Bayesian reconstruction of memories stored in neural networks from their connectivity
Goldt, Sebastian, Krzakala, Florent, Zdeborová, Lenka, Brunel, Nicolas
Comprehensive synaptic wiring diagrams or "connectomes" provide a detailed map of all the neurons and their interconnections in a brain region or even an entire organism. Since the connectome of the nematode C. elegans was obtained using electron microscopy methods in 1986 [1], methods for data acquisition and analysis have both been scaled up and improved significantly. Today, it has become possible to provide connectomes of much more complex systems such as various Drosophila melanogaster circuits [2, 3], or even a large part of its brain [4, 5]; the olfactory bulb of zebrafish [6]; and various pieces of the rodent retina [7-9], hippocampus [10], and cortex [11-14]. While there still remain a number of formidable challenges on the way to the complete connectome of a mammal or even human brain [15], the data sets available today already give rise to a number of intriguing questions. At the same time, it is becoming increasingly clear that new quantitative methods must be developed to fully exploit the new troves of data that connectomics provides [16]. Here, we focus on local neural networks that store information in their synaptic connectivity. It has been hypothesised that cortical networks with their extensive recurrent synaptic connectivity are optimised for this task [17]. A popular model for these networks are attractor neural networks such as the Hopfield's model [18] and various generalisations [19-22], where memories are stored as
Uncertainty in Minimum Cost Multicuts for Image and Motion Segmentation
Kardoost, Amirhossein, Keuper, Margret
The minimum cost lifted multicut approach has proven practically good performance in a wide range of applications such as image decomposition, mesh segmentation, multiple object tracking, and motion segmentation. It addresses such problems in a graph-based model, where real-valued costs are assigned to the edges between entities such that the minimum cut decomposes the graph into an optimal number of segments. Driven by a probabilistic formulation of minimum cost multicuts, we provide a measure for the uncertainties of the decisions made during the optimization. We argue that access to such uncertainties is crucial for many practical applications and conduct an evaluation by means of sparsifications on three different, widely used datasets in the context of image decomposition (BSDS-500) and motion segmentation (DAVIS2016 and FBMS59) in terms of variation of information (VI) and Rand index (RI).
Abstraction, Validation, and Generalization for Explainable Artificial Intelligence
Yang, Scott Cheng-Hsin, Folke, Tomas, Shafto, Patrick
Neural network architectures are achieving superhuman performance on an expanding range of tasks. To effectively and safely deploy these systems, their decision-making must be understandable to a wide range of stakeholders. Methods to explain AI have been proposed to answer this challenge, but a lack of theory impedes the development of systematic abstractions which are necessary for cumulative knowledge gains. We propose Bayesian Teaching as a framework for unifying explainable AI (XAI) by integrating machine learning and human learning. Bayesian Teaching formalizes explanation as a communication act of an explainer to shift the beliefs of an explainee. This formalization decomposes any XAI method into four components: (1) the inference to be explained, (2) the explanatory medium, (3) the explainee model, and (4) the explainer model. The abstraction afforded by Bayesian Teaching to decompose any XAI method elucidates the invariances among them. The decomposition of XAI systems enables modular validation, as each of the first three components listed can be tested semi-independently. This decomposition also promotes generalization through recombination of components from different XAI systems, which facilitates the generation of novel variants. These new variants need not be evaluated one by one provided that each component has been validated, leading to an exponential decrease in development time. Finally, by making the goal of explanation explicit, Bayesian Teaching helps developers to assess how suitable an XAI system is for its intended real-world use case. Thus, Bayesian Teaching provides a theoretical framework that encourages systematic, scientific investigation of XAI.
Order Effects in Bayesian Updates
Moreira, Catarina, de Barros, Jose Acacio
Order effects occur when judgments about a hypothesis's probability given a sequence of information do not equal the probability of the same hypothesis when the information is reversed. Different experiments have been performed in the literature that supports evidence of order effects. We proposed a Bayesian update model for order effects where each question can be thought of as a mini-experiment where the respondents reflect on their beliefs. We showed that order effects appear, and they have a simple cognitive explanation: the respondent's prior belief that two questions are correlated. The proposed Bayesian model allows us to make several predictions: (1) we found certain conditions on the priors that limit the existence of order effects; (2) we show that, for our model, the QQ equality is not necessarily satisfied (due to symmetry assumptions); and (3) the proposed Bayesian model has the advantage of possessing fewer parameters than its quantum counterpart.
CCMN: A General Framework for Learning with Class-Conditional Multi-Label Noise
Xie, Ming-Kun, Huang, Sheng-Jun
Class-conditional noise commonly exists in machine learning tasks, where the class label is corrupted with a probability depending on its ground-truth. Many research efforts have been made to improve the model robustness against the class-conditional noise. However, they typically focus on the single label case by assuming that only one label is corrupted. In real applications, an instance is usually associated with multiple labels, which could be corrupted simultaneously with their respective conditional probabilities. In this paper, we formalize this problem as a general framework of learning with Class-Conditional Multi-label Noise (CCMN for short). We establish two unbiased estimators with error bounds for solving the CCMN problems, and further prove that they are consistent with commonly used multi-label loss functions. Finally, a new method for partial multi-label learning is implemented with unbiased estimator under the CCMN framework. Empirical studies on multiple datasets and various evaluation metrics validate the effectiveness of the proposed method.
A causal learning framework for the analysis and interpretation of COVID-19 clinical data
Ferrari, Elisa, Gargani, Luna, Barbieri, Greta, Ghiadoni, Lorenzo, Faita, Francesco, Bacciu, Davide
We present a workflow for clinical data analysis that relies on Bayesian Structure Learning (BSL), an unsupervised learning approach, robust to noise and biases, that allows to incorporate prior medical knowledge into the learning process and that provides explainable results in the form of a graph showing the causal connections among the analyzed features. The workflow consists in a multi-step approach that goes from identifying the main causes of patient's outcome through BSL, to the realization of a tool suitable for clinical practice, based on a Binary Decision Tree (BDT), to recognize patients at high-risk with information available already at hospital admission time. We evaluate our approach on a feature-rich COVID-19 dataset, showing that the proposed framework provides a schematic overview of the multi-factorial processes that jointly contribute to the outcome. We discuss how these computational findings are confirmed by current understanding of the COVID-19 pathogenesis. Further, our approach yields to a highly interpretable tool correctly predicting the outcome of 85% of subjects based exclusively on 3 features: age, a previous history of chronic obstructive pulmonary disease and the PaO2/FiO2 ratio at the time of arrival to the hospital.