Statistical Learning
Reproduction Report on "Learn to Pay Attention"
Shugliashvili, Levan, Soselia, Davit, Amashukeli, Shota, Koberidze, Irakli
The model proposed in the "Learn to Pay Attention" paper introduced a novel way to generate a trainable attention module for convolutional neural networks. The paper demonstrated the attention modulein VGG-based and ResNet-based architectures, provided several options for implementation (includingthree options for layer depths at which attention modules are to be implemented, two options for calculating the compatibility between the global and local features in generating the attention maps, and two options for what method will be used to produce output probabilities from global-level feature vectors), described the dataset preprocessing and model training routines, and reported the results of the consequent models in several tasks. We have successfully implemented allpossible configurations of both VGG-based and ResNet-based attention models, and have replicated the paper's reported results in image classification and fine-grain recognition task using the (VGG-att2)-concat-pc configuration on the CIFAR-10 dataset and (VGG-att3)-concatpc configurationon CIFAR-100 and the SVHN dataset.
Kernel Treelets
Xia, Hedi, Ceniceros, Hector D.
Treelets, introduced by Lee, Nadler, and Wasserman [1, 2], is a method to produce a multiscale, hierarchicaldecomposition of unordered data. The central premise of Treelets is to exploit sparsity and capture intrinsic localized structures with only a few features, represented interms of an orthonormal basis. The hierarchical tree constructed by the treelet algorithm provides a scale-based partition of the data that can be used for classification, specially for cluster analysis [3]. Cluster analysis, also called clustering, is concerned with finding a partition of a set such that its corresponding equivalence class captures similarity of its elements. The Treelet approach is an example of hierarchical clustering (HC) [4], which is a type of methods that provides a nested and multiscale clustering.
Towards Automatic Personality Prediction Using Facebook Like Categories
Tareaf, Raad Bin, Berger, Philipp, Hennig, Patrick, Meinel, Christoph
We demonstrate that effortlessly accessible digital records of behavior such as Facebook Likes can be obtained and utilized to automatically distinguish a wide range of highly delicate personal traits including: life satisfaction, cultural ethnicity, political views, age, gender and personality traits. The analysis presented based on a dataset of over 738,000 users who conferred their Facebook Likes, social network activities, egocentric network, demographic characteristics, and the results of various psychometric tests for our extended personality analysis. The proposed model uses unique mapping technique between each Facebook Like object to the corresponding Facebook page category/sub-category object, which is then evaluated as features for a set of machine learning algorithms to predict individual psycho-demographic profiles from Likes. The model , distinguishes between a religious and non-religious individual in 83% of circumstances, Asian and European in 87% of situations, and between emotional stable and emotion unstable in 81% of situations. We provide exemplars of correlations between attributes and Likes and present suggestions for future directions.
From Adaptive Kernel Density Estimation to Sparse Mixture Models
Schretter, Colas, Sun, Jianyong, Schelkens, Peter
We introduce a balloon estimator in a generalized expectation-maximization method for estimating all parameters of a Gaussian mixture model given one data sample per mixture component. Instead of limiting explicitly the model size, this regularization strategy yields low-complexity sparse models where the number of effective mixture components reduces with an increase of a smoothing probability parameter $\mathbf{P>0}$. This semi-parametric method bridges from non-parametric adaptive kernel density estimation (KDE) to parametric ordinary least-squares when $\mathbf{P=1}$. Experiments show that simpler sparse mixture models retain the level of details present in the adaptive KDE solution.
Online Newton Step Algorithm with Estimated Gradient
Liu, Binbin, Li, Jundong, Song, Yunquan, Liang, Xijun, Jian, Ling, Liu, Huan
They have shown to be effective in handling large-scale and high-velocity streaming data and emerged to become popular in the big data era Hoi et al. [2018, 2014]. In recent years, a number of effective online learning algorithms have been investigated and applied in a variety of high impact domains,ranging from game theory, information theory to machine learning and data mining Ding et al. [2017], Shalev-Shwartz [2011], Wang et al. [2003]. Most previously proposed online learning algorithms fall into the wellestablished frameworkof online convex optimization Gordon [1999], Zinkevich [2003]. In terms of the optimization algorithms, online learning algorithms can be grouped into the following categories: (i) first-order algorithms which aim to optimize the objective function using the first-order (sub) gradient such as the well-known OGD algorithm Zinkevich [2003]; and (ii) second-order algorithms which aim to exploit second-order information to speed up the convergence of the optimization, such as the ONS algorithm Hazan et al. [2007]. In online convex optimization, previous approaches are mainly based on the first-order optimization, i.e., optimization using the first-order derivative of the cost function. Theregret bound achieved by these algorithms is proportional to the polynomial of the number of rounds T . For example, Zinkevich [2003] showed that with the simple OGD, we can achieve the regret bound of O( T). Later on, Hazan et al. [2007] introduced a new algorithm with ONS by exploiting the second-order derivative of the cost function, which can be viewed as an online 2 analogy of the Newton-Raphson method Ypma and Tjalling [1995] in the offline learning.Although the time complexity O(d
Robust Bregman Clustering
Brรฉcheteau, Claire, Fischer, Aurรฉlie, Levrard, Clรฉment
Using a trimming approach, we investigate a k-means type method based on Bregman divergences for clustering data possibly corrupted with clutter noise. The main interest of Bregman divergences is that the standard Lloyd algorithm adapts to these distortion measures, and they are well-suited for clustering data sampled according to mixture models from exponential families. We prove that there exists an optimal codebook, and that an empirically optimal codebook converges a.s. to an optimal codebook in the distortion sense. Moreover, we obtain the sub-Gaussian rate of convergence for k-means 1 $\sqrt$ n under mild tail assumptions. Also, we derive a Lloyd-type algorithm with a trimming parameter that can be selected from data according to some heuristic, and present some experimental results.
The FLUXCOM ensemble of global land-atmosphere energy fluxes
Jung, Martin, Koirala, Sujan, Weber, Ulrich, Ichii, Kazuhito, Gans, Fabian, Gustau-Camps-Valls, null, Papale, Dario, Schwalm, Christopher, Tramontana, Gianluca, Reichstein, Markus
Although a key driver of Earth's climate system, global land-atmosphere energy fluxes are poorly constrained. Here we use machine learning to merge energy flux measurements from FLUXNET eddy covariance towers with remote sensing and meteorological data to estimate net radiation, latent and sensible heat and their uncertainties. The resulting FLUXCOM database comprises 147 global gridded products in two setups: (1) 0.0833${\deg}$ resolution using MODIS remote sensing data (RS) and (2) 0.5${\deg}$ resolution using remote sensing and meteorological data (RS+METEO). Within each setup we use a full factorial design across machine learning methods, forcing datasets and energy balance closure corrections. For RS and RS+METEO setups respectively, we estimate 2001-2013 global (${\pm}$ 1 standard deviation) net radiation as 75.8${\pm}$1.4 ${W\ m^{-2}}$ and 77.6${\pm}$2 ${W\ m^{-2}}$, sensible heat as 33${\pm}$4 ${W\ m^{-2}}$ and 36${\pm}$5 ${W\ m^{-2}}$, and evapotranspiration as 75.6${\pm}$10 ${\times}$ 10$^3$ ${km^3\ yr^{-1}}$ and 76${\pm}$6 ${\times}$ 10$^3$ ${km^3\ yr^{-1}}$. FLUXCOM products are suitable to quantify global land-atmosphere interactions and benchmark land surface model simulations.
Semi-supervised dual graph regularized dictionary learning
Tran, Khanh-Hung, Ngole-Mboula, Fred-Maurice, Starck, Jean-Luc
Dictionary Learning (DL) encompasses methods and algorithms thataim at deriving a set of cardinal features which enables one to concisely describe signals of a given type. The benefit of such dictionaries in sparsity-driven signal recovery hasbeen shown in several applications (see for example [1, 2, 3, 4]). In numerous applications of machine learning, data are labelled and/orsampled from some regular manifold; thus it is suitable, for classification or interpolation tasks for instance, that the learned codes allow for a better discrimination of the data samples with respect to labels information or manifold's structure. The growing field of supervised dictionary learning precisely consists of DL methods that account for these additional information(a recent review can be found in [5]). Unlike the supervised classification in which only labelled data is used to train the classifier, the unlabelled data is also used in training to make use of all the manifold's structure information available.
Deep Density-based Image Clustering
Ren, Yazhou, Wang, Ni, Li, Mingxia, Xu, Zenglin
Recently, deep clustering, which is able to perform feature learning that favors clustering tasks via deep neural networks, has achieved remarkable performance in image clustering applications. However, the existing deep clustering algorithms generally need the number of clusters in advance, which is usually unknown in real-world tasks. In addition, the initial cluster centers in the learned feature space are generated by $k$-means. This only works well on spherical clusters and probably leads to unstable clustering results. In this paper, we propose a two-stage deep density-based image clustering (DDC) framework to address these issues. The first stage is to train a deep convolutional autoencoder (CAE) to extract low-dimensional feature representations from high-dimensional image data, and then apply t-SNE to further reduce the data to a 2-dimensional space favoring density-based clustering algorithms. The second stage is to apply the developed density-based clustering technique on the 2-dimensional embedded data to automatically recognize an appropriate number of clusters with arbitrary shapes. Concretely, a number of local clusters are generated to capture the local structures of clusters, and then are merged via their density relationship to form the final clustering result. Experiments demonstrate that the proposed DDC achieves comparable or even better clustering performance than state-of-the-art deep clustering methods, even though the number of clusters is not given.
Variational Bayesian Complex Network Reconstruction
Xu, Shuang, Zhang, Chun-Xia, Wang, Pei, Zhang, Jiangshe
The networked systems are ubiquitous in many fields, including social-tech science [1, 2], bioinformatics [3-6], epidemic dynamics [7-9] and power grid [10, 11]. However, as is often the case, it is not able to observe the topology of a network, while data generated by this network are available. Therefore, in interdisciplinary science, one of the most important but challenging problems is to reconstruct the complex network from the observed data or time series [12]. This problem has been widely investigated in the past three decades, where the classical method is the delay-coordinate embedding method proposed by Takens [13], which, nevertheless, is only suitable for small-scale networks. Nowadays, with the advent of big data era [14], it is of great urgency solve this issue for large-scale complex networks. Suppose that a complex network consists of N nodes, in practice we are often given the time series of the states for the N nodes. Generally speaking, the core idea of many data-driven network reconstruction investigations is to first calculate the correlation between two nodes. Then, a threshold can be set mutually or automatically to make the network binary.