Goto

Collaborating Authors

 Statistical Learning


Just a Momentum: Analytical Study of Momentum-Based Acceleration Methods in Paradigmatic High-Dimensional Non-Convex Problems

arXiv.org Machine Learning

When optimizing over loss functions it is common practice to use momentum-based accelerated methods rather than vanilla gradient-based method. Despite widely applied to arbitrary loss function, their behaviour in generically non-convex, high dimensional landscapes is poorly understood. In this work we used dynamical mean field theory techniques to describe analytically the average behaviour of these methods in a prototypical non-convex model: the (spiked) matrix-tensor model. We derive a closed set of equations that describe the behaviours of several algorithms including heavy-ball momentum and Nesterov acceleration. Additionally we characterize the evolution of a mathematically equivalent physical system of massive particles relaxing toward the bottom of an energetic landscape. Under the correct mapping the two dynamics are equivalent and it can be noticed that having a large mass increases the effective time step of the heavy ball dynamics leading to a speed up.


DynACPD Embedding Algorithm for Prediction Tasks in Dynamic Networks

arXiv.org Artificial Intelligence

Classical network embeddings create a low dimensional representation of the learned relationships between features across nodes. Such embeddings are important for tasks such as link prediction and node classification. In the current paper, we consider low dimensional embeddings of dynamic networks, that is a family of time varying networks where there exist both temporal and spatial link relationships between nodes. We present novel embedding methods for a dynamic network based on higher order tensor decompositions for tensorial representations of the dynamic network. In one sense, our embeddings are analogous to spectral embedding methods for static networks. We provide a rationale for our algorithms via a mathematical analysis of some potential reasons for their effectiveness. Finally, we demonstrate the power and efficiency of our approach by comparing our algorithms' performance on the link prediction task against an array of current baseline methods across three distinct real-world dynamic networks.


Integrated Age Estimation Mechanism

arXiv.org Artificial Intelligence

Machine-learning-based age estimation has received lots of attention. Traditional age estimation mechanism focuses estimation age error, but ignores that there is a deviation between the estimated age and real age due to disease. Pathological age estimation mechanism the author proposed before introduces age deviation to solve the above problem and improves classification capability of the estimated age significantly. However,it does not consider the age estimation error of the normal control (NC) group and results in a larger error between the estimated age and real age of NC group. Therefore, an integrated age estimation mechanism based on Decision-Level fusion of error and deviation orientation model is proposed to solve the problem.Firstly, the traditional age estimation and pathological age estimation mechanisms are weighted together.Secondly, their optimal weights are obtained by minimizing mean absolute error (MAE) between the estimated age and real age of normal people. In the experimental section, several representative age-related datasets are used for verification of the proposed method. The results show that the proposed age estimation mechanism achieves a good tradeoff effect of age estimation. It not only improves the classification ability of the estimated age, but also reduces the age estimation error of the NC group. In general, the proposed age estimation mechanism is effective. Additionally, the mechanism is a framework mechanism that can be used to construct different specific age estimation algorithms, contributing to relevant research.


Metapaths guided Neighbors aggregated Network for?Heterogeneous Graph Reasoning

arXiv.org Artificial Intelligence

Most real-world datasets are inherently heterogeneous graphs, which involve a diversity of node and relation types. Heterogeneous graph embedding is to learn the structure and semantic information from the graph, and then embed it into the low-dimensional node representation. Existing methods usually capture the composite relation of a heterogeneous graph by defining metapath, which represent a semantic of the graph. However, these methods either ignore node attributes, or discard the local and global information of the graph, or only consider one metapath. To address these limitations, we propose a Metapaths-guided Neighbors-aggregated Heterogeneous Graph Neural Network(MHN) to improve performance. Specially, MHN employs node base embedding to encapsulate node attributes, BFS and DFS neighbors aggregation within a metapath to capture local and global information, and metapaths aggregation to combine different semantics of the heterogeneous graph. We conduct extensive experiments for the proposed MHN on three real-world heterogeneous graph datasets, including node classification, link prediction and online A/B test on Alibaba mobile application. Results demonstrate that MHN performs better than other state-of-the-art baselines.


AutoDO: Robust AutoAugment for Biased Data with Label Noise via Scalable Probabilistic Implicit Differentiation

arXiv.org Artificial Intelligence

AutoAugment has sparked an interest in automated augmentation methods for deep learning models. These methods estimate image transformation policies for train data that improve generalization to test data. While recent papers evolved in the direction of decreasing policy search complexity, we show that those methods are not robust when applied to biased and noisy data. To overcome these limitations, we reformulate AutoAugment as a generalized automated dataset optimization (AutoDO) task that minimizes the distribution shift between test data and distorted train dataset. In our AutoDO model, we explicitly estimate a set of per-point hyperparameters to flexibly change distribution of train data. In particular, we include hyperparameters for augmentation, loss weights, and soft-labels that are jointly estimated using implicit differentiation. We develop a theoretical probabilistic interpretation of this framework using Fisher information and show that its complexity scales linearly with the dataset size. Our experiments on SVHN, CIFAR-10/100, and ImageNet classification show up to 9.3% improvement for biased datasets with label noise compared to prior methods and, importantly, up to 36.6% gain for underrepresented SVHN classes.


Gradient Descent vs Stochastic GD vs Mini-Batch SGD

#artificialintelligence

Warning: Just in case the terms "partial derivative" or "gradient" sound unfamiliar, I suggest checking out these resources! Gradient descent is an iterative algorithm whose purpose is to make changes to a set of parameters (i.e. A loss or cost or objective function (any of these naming conventions work in practice) is the function whose value we seek to minimize. When performing Gradient descent, each time we update the parameters, we expect to observe a change in min f(w). That is at each iteration, the gradient of the function that contains parameters in w is taken so that changes in the function with respect to parameters brings us closer to the goal of reaching an optimal set of parameters that will ultimately lead to the lowest possible loss function value.


Machine Learning Adds Little to MI Prognostication

#artificialintelligence

Machine learning (ML) algorithms developed to predict in-hospital mortality after acute MI offered more meaningful gains in model calibration than in accuracy, researchers found. Parsing through data on 29 variables from the American College of Cardiology (ACC) Chest Pain-MI Registry, extreme gradient descent boosting (XGBoost) and meta-classifier models offered no substantive improvement in discrimination compared with standard logistic regression modeling (C-statistics 0.90 for both vs 0.89), reported Harlan Krumholz, MD, SM, of Yale School of Medicine, and colleagues. However, the two ML models showed nearly perfect agreement between observed and predicted risk across the risk spectrum. Of the people deemed moderate-to-high risk in logistic regression, 27% were more accurately reclassified as low risk by the XGBoost model and 25% by the meta-classifier model -- both more consistent with the observed event rates. "These findings suggest that ML models are not associated with substantially better prediction of risk of death after acute MI but may offer greater resolution of risk, which can better clarify the individual risk for adverse outcomes," Krumholz's group reported in a paper published online in JAMA Cardiology.


K-Nearest Neighbor - 360DigitMG

#artificialintelligence

An Artificial Neural Network (ANN) models the relationship between a set of input signals and an output signal using a model derived from our understanding of how a biological brain responds to stimuli from sensory inputs.


Optimal Targeting in Fundraising: A Machine Learning Approach

arXiv.org Machine Learning

Fundraising is a costly activity: the largest 25 US charities spend between 5% and 25% of total donations on fundraising expenses (Andreoni and Payne, 2011). These numbers are a matter of concern for two reasons. First, high fundraising costs imply that a smaller proportion of overall donations can finance charitable projects. This effect can lead to an underprovision of the provided goods and services and may, thus, lower welfare if the donors' utility depends on provision levels (Rose-Ackerman, 1982; Name-Correa and Yildirim, 2013). Second, high fundraising costs also matter from the charities' perspectives: it is well documented that donors are averse to financing overhead costs (Tinkelman and Mankaney, 2007; Gneezy et al., 2014). Hence, charities with excessive fundraising expenses will be less successful in raising donations. In conclusion, reducing disproportional fundraising costs can be crucial, both from a welfare and a charity-management perspective. However, while there is a broad literature studying how fundraising instruments such as matching grants and unconditional gifts affect donors' behavior (surveyed by Andreoni and Payne, 2013), previous research has paid less attention to how charities could increase the cost efficacy of fundraising. This paper shifts focus to a novel approach to increase a fundraising campaigns' efficacy: optimal targeting of fundraising activities based on causal machine learning.


Covariate-assisted Sparse Tensor Completion

arXiv.org Machine Learning

We aim to provably complete a sparse and highly-missing tensor in the presence of covariate information along tensor modes. Our motivation comes from online advertising where users click-through-rates (CTR) on ads over various devices form a CTR tensor that has about 96% missing entries and has many zeros on non-missing entries, which makes the standalone tensor completion method unsatisfactory. Beside the CTR tensor, additional ad features or user characteristics are often available. In this paper, we propose Covariate-assisted Sparse Tensor Completion (COSTCO) to incorporate covariate information for the recovery of the sparse tensor. The key idea is to jointly extract latent components from both the tensor and the covariate matrix to learn a synthetic representation. Theoretically, we derive the error bound for the recovered tensor components and explicitly quantify the improvements on both the reveal probability condition and the tensor recovery accuracy due to covariates. Finally, we apply COSTCO to an advertisement dataset consisting of a CTR tensor and ad covariate matrix, leading to 23% accuracy improvement over the baseline. An important by-product is that ad latent components from COSTCO reveal interesting ad clusters, which are useful for better ad targeting.