Genre
Approximate maximum entropy principles via Goemans-Williamson with applications to provable variational methods
The well known maximum-entropy principle due to Jaynes, which states that given mean parameters, the maximum entropy distribution matching them is in an exponential family, has been very popular in machine learning due to its "Occam's razor" interpretation. Unfortunately, calculating the potentials in the maximum-entropy distribution is intractable \cite{bresler2014hardness}. We provide computationally efficient versions of this principle when the mean parameters are pairwise moments: we design distributions that approximately match given pairwise moments, while having entropy which is comparable to the maximum entropy distribution matching those moments. We additionally provide surprising applications of the approximate maximum entropy principle to designing provable variational methods for partition function calculations for Ising models without any assumptions on the potentials of the model. More precisely, we show that in every temperature, we can get approximation guarantees for the log-partition function comparable to those in the low-temperature limit, which is the setting of optimization of quadratic forms over the hypercube. \cite{alon2006approximating}
Predicting the evolution of stationary graph signals
Loukas, Andreas, Perraudin, Nathanael
In the problem of modeling and predicting statistical processes, (wide-sense) stationarity is a helpful assumption, that allows us to learn the spectral characteristics of a process using very few samples. Especially for time-series prediction, learning from few samples is crucial, as one needs to estimate future values after only partially observing a single realization of the statistical process. This is the main reason why classical models for estimation and prediction of univariate processes, such as Wiener filters and auto-regressive moving average models (ARMA), rely on stationarity to produce predictions. For multivariate statistical processes, following the same methodology is often problematic, as the number of parameters to be estimated increases quadratically with the number variables, often rendering the problem intractable. A common way to deal with this dimensionality issue is to assume that there is an inherent structure to the process that can be captured by a graph. The graph assumption appears frequently in the machine learning and signal processing literature, and has been shown invaluable for tasks such as clustering [1], [18], low-rank extraction [14], spectral estimation [12], [10] and semi-supervised learning [2], [17]. Nevertheless, despite their promise, so far state-of-the-art graph-based methods predominantly ignore the time-dimension of data. The objective of this paper is to identify multivariate models that exploit the graph structure inherent to the data so as to facilitate the task of prediction.
Combining multiple resolutions into hierarchical representations for kernel-based image classification
Cui, Yanwei, Lefevre, Sรฉbastien, Chapel, Laetitia, Puissant, Anne
Geographic object-based image analysis (GEOBIA) framework has gained increasing interest recently. Following this popular paradigm, we propose a novel multiscale classification approach operating on a hierarchical image representation built from two images at different resolutions. They capture the same scene with different sensors and are naturally fused together through the hierarchical representation, where coarser levels are built from a Low Spatial Resolution (LSR) or Medium Spatial Resolution (MSR) image while finer levels are generated from a High Spatial Resolution (HSR) or Very High Spatial Resolution (VHSR) image. Such a representation allows one to benefit from the context information thanks to the coarser levels, and subregions spatial arrangement information thanks to the finer levels. Two dedicated structured kernels are then used to perform machine learning directly on the constructed hierarchical representation. This strategy overcomes the limits of conventional GEOBIA classification procedures that can handle only one or very few pre-selected scales. Experiments run on an urban classification task show that the proposed approach can highly improve the classification accuracy w.r.t.
UBL: an R package for Utility-based Learning
Branco, Paula, Ribeiro, Rita P., Torgo, Luis
This document describes the R package UBL that allows the use of several methods for handling utility-based learning problems. Classification and regression problems that assume non-uniform costs and/or benefits pose serious challenges to predictive analytic tasks. In the context of meteorology, finance, medicine, ecology, among many other, specific domain information concerning the preference bias of the users must be taken into account to enhance the models predictive performance. To deal with this problem, a large number of techniques was proposed by the research community for both classification and regression tasks. The main goal of UBL package is to facilitate the utility-based predictive analytic task by providing a set of methods to deal with this type of problems in the R environment. It is a versatile tool that provides mechanisms to handle both regression and classification (binary and multiclass) tasks. Moreover, UBL package allows the user to specify his domain preferences, but it also provides some automatic methods that try to infer those preference bias from the domain, considering some common known settings.
Normalization Propagation: A Parametric Technique for Removing Internal Covariate Shift in Deep Networks
Arpit, Devansh, Zhou, Yingbo, Kota, Bhargava U., Govindaraju, Venu
While the authors of Batch Normalization (BN) identify and address an important problem involved in training deep networks-- Internal Covariate Shift-- the current solution has certain drawbacks. Specifically, BN depends on batch statistics for layerwise input normalization during training which makes the estimates of mean and standard deviation of input (distribution) to hidden layers inaccurate for validation due to shifting parameter values (especially during initial training epochs). Also, BN cannot be used with batch-size 1 during training. We address these drawbacks by proposing a non-adaptive normalization technique for removing internal covariate shift, that we call Normalization Propagation. Our approach does not depend on batch statistics, but rather uses a data-independent parametric estimate of mean and standard-deviation in every layer thus being computationally faster compared with BN. We exploit the observation that the pre-activation before Rectified Linear Units follow Gaussian distribution in deep networks, and that once the first and second order statistics of any given dataset are normalized, we can forward propagate this normalization without the need for recalculating the approximate statistics for hidden layers.
Local identifiability of $l_1$-minimization dictionary learning: a sufficient and almost necessary condition
We study the theoretical properties of learning a dictionary from $N$ signals $\mathbf x_i\in \mathbb R^K$ for $i=1,...,N$ via $l_1$-minimization. We assume that $\mathbf x_i$'s are $i.i.d.$ random linear combinations of the $K$ columns from a complete (i.e., square and invertible) reference dictionary $\mathbf D_0 \in \mathbb R^{K\times K}$. Here, the random linear coefficients are generated from either the $s$-sparse Gaussian model or the Bernoulli-Gaussian model. First, for the population case, we establish a sufficient and almost necessary condition for the reference dictionary $\mathbf D_0$ to be locally identifiable, i.e., a local minimum of the expected $l_1$-norm objective function. Our condition covers both sparse and dense cases of the random linear coefficients and significantly improves the sufficient condition by Gribonval and Schnass (2010). In addition, we show that for a complete $\mu$-coherent reference dictionary, i.e., a dictionary with absolute pairwise column inner-product at most $\mu\in[0,1)$, local identifiability holds even when the random linear coefficient vector has up to $O(\mu^{-2})$ nonzeros on average. Moreover, our local identifiability results also translate to the finite sample case with high probability provided that the number of signals $N$ scales as $O(K\log K)$.
Natural brain-information interfaces: Recommending information by relevance inferred from human brain signals
Eugster, Manuel J. A., Ruotsalo, Tuukka, Spapรฉ, Michiel M., Barral, Oswald, Ravaja, Niklas, Jacucci, Giulio, Kaski, Samuel
Finding relevant information from large document collections such as the World Wide Web is a common task in our daily lives. Estimation of a user's interest or search intention is necessary to recommend and retrieve relevant information from these collections. We introduce a brain-information interface used for recommending information by relevance inferred directly from brain signals. In experiments, participants were asked to read Wikipedia documents about a selection of topics while their EEG was recorded. Based on the prediction of word relevance, the individual's search intent was modeled and successfully used for retrieving new, relevant documents from the whole English Wikipedia corpus. The results show that the users' interests towards digital content can be modeled from the brain signals evoked by reading. The introduced brain-relevance paradigm enables the recommendation of information without any explicit user interaction, and may be applied across diverse information-intensive applications.
Top 10 Machine Learning Algorithms
According to a recent study, machine learning algorithms are expected to replace 25% of the jobs across the world, in the next 10 years. With the rapid growth of big data and availability of programming tools like Python and R โmachine learning is gaining mainstream presence for data scientists. Machine learning applications are highly automated and self-modifying which continue to improve over time with minimal human intervention as they learn with more data. For instance, Netflix's recommendation algorithm learns more about the likes and dislikes of a viewer based on the shows every viewer watches. To address the complex nature of various real world data problems, specialized machine learning algorithms have been developed that solve these problems perfectly.
Microsoft, GE partner to link industrial machines with cloud
Microsoft CEO Satya Nadella, at left, with Jeff Immelt, CEO of GE. (Photo: Brian Smale) Microsoft and GE are joining forces to use cloud technologies in industrial production. GE will bring its Predix industrial data-gathering operating system to Microsoft's Azure cloud computing platform, the two companies announced Monday at the Microsoft Worldwide Partner Conference in Toronto. The deal is the first in a planned strategic partnership between GE and Microsoft. "Every industry and every company around the world is being transformed by digital technology," said Microsoft CEO Satya Nadella in a statement announcing the partnership. "Working with companies like GE, we can reach a new set of customers to help them accelerate their transformation across every line of business -- from the factory floor to smart buildings."
Report: Machine Learning Driving AI
Artificial intelligence research continues to accelerate as human and machines collaborate to solve more complex problems. A new survey by the National Academy of Science identifies the frontiers of AI research that include "augmented cognition" along with "integrative" AI. In a workshop report on IT innovation just released by the National Academy's science, engineering and medicine branches, a section on "Developing Smart Machines" describes ongoing AI research efforts along with machine learning, a discipline some experts consider a subfield of AI. "Machine learning is what lets computers discover patterns within data and then use those patterns to make useful, and ideally correct, predictions," the report states. The authors quotes Jaime Carbonell, a professor of computer science at Carnegie Mellon University, as noting" "Machine learning essentially is the engine that is driving modern artificial intelligence." Machine learning "often deals with unbalanced data sets in which the ultimate focus of decision making is precisely the outlier cases," that is, the extreme cases that contrast sharply with typical data and provide opportunities for the "most important learning opportunities."