Goto

Collaborating Authors

 Statistical Learning


Self-explaining variational posterior distributions for Gaussian Process models

arXiv.org Machine Learning

Bayesian methods have become a popular way to incorporate prior knowledge and a notion of uncertainty into machine learning models. At the same time, the complexity of modern machine learning makes it challenging to comprehend a model's reasoning process, let alone express specific prior assumptions in a rigorous manner. While primarily interested in the former issue, recent developments in transparent machine learning could also broaden the range of prior information that we can provide to complex Bayesian models. Inspired by the idea of selfexplaining models, we introduce a corresponding concept for variational Gaussian Processes. On the one hand, our contribution improves transparency for these types of models. More importantly though, our proposed self-explaining variational posterior distribution allows to incorporate both general prior knowledge about a target function as a whole and prior knowledge about the contribution of individual features.


Higher Order Kernel Mean Embeddings to Capture Filtrations of Stochastic Processes

arXiv.org Machine Learning

Stochastic processes are random variables with values in some space of paths. However, reducing a stochastic process to a path-valued random variable ignores its filtration, i.e. the flow of information carried by the process through time. By conditioning the process on its filtration, we introduce a family of higher order kernel mean embeddings (KMEs) that generalizes the notion of KME and captures additional information related to the filtration. We derive empirical estimators for the associated higher order maximum mean discrepancies (MMDs) and prove consistency. We then construct a filtration-sensitive kernel two-sample test able to pick up information that gets missed by the standard MMD test. In addition, leveraging our higher order MMDs we construct a family of universal kernels on stochastic processes that allows to solve real-world calibration and optimal stopping problems in quantitative finance (such as the pricing of American options) via classical kernel-based regression methods. Finally, adapting existing tests for conditional independence to the case of stochastic processes, we design a causal-discovery algorithm to recover the causal graph of structural dependencies among interacting bodies solely from observations of their multidimensional trajectories.


Explained: Linear Regression with real life scenarios in R

#artificialintelligence

Machine learning is one of the most trending topics at present and is expected to grow exponentially over the coming years. Before we drill down to one of the most common techniques in machine learning that is the'linear regression', let's understand what exactly is regression. Regression analysis is a form of a predictive modelling technique that establishes a relationship between two variables namely a dependent variable and independent variable. In simpler words, a regression analysis involves graphing a line over a set of data points that most closely fits the overall shape of the data or it can be said that a regression shows the changes in a dependent variable on the y-axis to the change in the explanatory variable on the x-axis. Linear regression aims to establish a linear relationship between two variables where one is the independent variable and other is the dependent variable using a linear equation on the data that is put under observation. For example if we consider the weight and height of a person, as the height increases the weight also increases, hence a linear relationship can be established between the height and weight of a person.


An interaction regression model for crop yield prediction - Scientific Reports

#artificialintelligence

Crop yield prediction is crucial for global food security yet notoriously challenging due to multitudinous factors that jointly determine the yield, including genotype, environment, management, and their complex interactions. Integrating the power of optimization, machine learning, and agronomic insight, we present a new predictive model (referred to as the interaction regression model) for crop yield prediction, which has three salient properties. First, it achieved a relative root mean square error of 8% or less in three Midwest states (Illinois, Indiana, and Iowa) in the US for both corn and soybean yield prediction, outperforming state-of-the-art machine learning algorithms. Second, it identified about a dozen environment by management interactions for corn and soybean yield, some of which are consistent with conventional agronomic knowledge whereas some others interactions require additional analysis or experiment to prove or disprove. Third, it quantitatively dissected crop yield into contributions from weather, soil, management, and their interactions, allowing agronomists to pinpoint the factors that favorably or unfavorably affect the yield of a given location under a given weather and management scenario. The most significant contribution of the new prediction model is its capability to produce accurate prediction and explainable insights simultaneously. This was achieved by training the algorithm to select features and interactions that are spatially and temporally robust to balance prediction accuracy for the training data and generalizability to the test data.


Predicting the Cellular Localization Sites of Proteins in Yest - Projects Based Learning

#artificialintelligence

Convert String data to Numeric format so we can process the data in Apache Spark ML Library. Welcome to this project on predicting the Cellular Localization Sites of Proteins in Yest in Apache Spark Machine Learning using Databricks platform community edition server which allows you to execute your spark code, free of cost on their server just by registering through email id. In this project, we explore Apache Spark and Machine Learning on the Databricks platform. I am a firm believer that the best way to learn is by doing. That's why I haven't included any purely theoretical lectures in this tutorial: you will learn everything on the way and be able to put it into practice straight away.


Adaptive variational Bayes: Optimality, computation and applications

arXiv.org Machine Learning

In this paper, we explore adaptive inference based on variational Bayes. Although a number of studies have been conducted to analyze contraction properties of variational posteriors, there is still a lack of a general and computationally tractable variational Bayes method that can achieve adaptive optimal contraction of the variational posterior. We propose a novel variational Bayes framework, called adaptive variational Bayes, which can operate on a collection of models with varying dimensions and structures. The proposed framework combines variational posteriors over individual models with certain weights to obtain a variational posterior over the entire model. It turns out that this combined variational posterior minimizes the Kullback-Leibler divergence to the original posterior distribution. We show that the proposed variational posterior achieves optimal contraction rates adaptively under very general conditions and attains model selection consistency when the true model structure exists. We apply the general results obtained for the adaptive variational Bayes to several examples including deep learning models and derive some new and adaptive inference results. Moreover, we consider the use of quasi-likelihood in our framework. We formulate conditions on the quasi-likelihood to ensure the adaptive optimality and discuss specific applications to stochastic block models and nonparametric regression with sub-Gaussian errors.


Biologically Plausible Learning Rules for Perceptual Systems that Maximize Mutual Information

arXiv.org Artificial Intelligence

Consider a neural perceptual system being exposed to an external environment. The system has certain internal state to represent external events. There is strong behavioral and neural evidence(e.g., Ernst and Banks, 2002; Gabbiani and Koch, 1998) that the internal representation is intrinsically probabilistic(Knill and Pouget, 2004), in line with the statistical properties of the environment. We mark the input signal as x. The perceptual representation would be a probability distribution conditional on x, denoted as p(y x). According to the Infomax principle (Attneave, 1954; Barlow et al., 1961; Linsker, 1988), the system's goal is to maximize the mutual information (MI) between the input x and the output (neuronal response) y, which can be written as max I(x;y), (1.1)


Learning grounded word meaning representations on similarity graphs

arXiv.org Artificial Intelligence

This paper introduces a novel approach to learn visually grounded meaning representations of words as low-dimensional node embeddings on an underlying graph hierarchy. The lower level of the hierarchy models modality-specific word representations through dedicated but communicating graphs, while the higher level puts these representations together on a single graph to learn a representation jointly from both modalities. The topology of each graph models similarity relations among words, and is estimated jointly with the graph embedding. The assumption underlying this model is that words sharing similar meaning correspond to communities in an underlying similarity graph in a low-dimensional space. We named this model Hierarchical Multi-Modal Similarity Graph Embedding (HM-SGE). Experimental results validate the ability of HM-SGE to simulate human similarity judgements and concept categorization, outperforming the state of the art.


Naturalness Evaluation of Natural Language Generation in Task-oriented Dialogues using BERT

arXiv.org Artificial Intelligence

This paper presents an automatic method to evaluate the naturalness of natural language generation in dialogue systems. While this task was previously rendered through expensive and time-consuming human labor, we present this novel task of automatic naturalness evaluation of generated language. By fine-tuning the BERT model, our proposed naturalness evaluation method shows robust results and outperforms the baselines: support vector machines, bi-directional LSTMs, and BLEURT. In addition, the training speed and evaluation performance of naturalness model are improved by transfer learning from quality and informativeness linguistic knowledge.


Fishr: Invariant Gradient Variances for Out-of-distribution Generalization

arXiv.org Artificial Intelligence

Learning robust models that generalize well under changes in the data distribution is critical for real-world applications. To this end, there has been a growing surge of interest to learn simultaneously from multiple training domains - while enforcing different types of invariance across those domains. Yet, all existing approaches fail to show systematic benefits under fair evaluation protocols. In this paper, we propose a new learning scheme to enforce domain invariance in the space of the gradients of the loss function: specifically, we introduce a regularization term that matches the domain-level variances of gradients across training domains. Critically, our strategy, named Fishr, exhibits close relations with the Fisher Information and the Hessian of the loss. We show that forcing domain-level gradient covariances to be similar during the learning procedure eventually aligns the domain-level loss landscapes locally around the final weights. Extensive experiments demonstrate the effectiveness of Fishr for out-of-distribution generalization. In particular, Fishr improves the state of the art on the DomainBed benchmark and performs significantly better than Empirical Risk Minimization. The code is released at https://github.com/alexrame/fishr.