Genre
Kernel regression, minimax rates and effective dimensionality: beyond the regular case
Blanchard, Gilles, Mücke, Nicole
We investigate if kernel regularization methods can achieve minimax convergence rates over a source condition regularity assumption for the target function. These questions have been considered in past literature, but only under specific assumptions about the decay, typically polynomial, of the spectrum of the the kernel mapping covariance operator. In the perspective of distribution-free results, we investigate this issue under much weaker assumption on the eigenvalue decay, allowing for more complex behavior that can reflect different structure of the data at different scales.
GANS for Sequences of Discrete Elements with the Gumbel-softmax Distribution
Kusner, Matt J., Hernández-Lobato, José Miguel
Generative Adversarial Networks (GAN) have limitations when the goal is to generate sequences of discrete elements. The reason for this is that samples from a distribution on discrete objects such as the multinomial are not differentiable with respect to the distribution parameters. This problem can be avoided by using the Gumbel-softmax distribution, which is a continuous approximation to a multinomial distribution parameterized in terms of the softmax function. In this work, we evaluate the performance of GANs based on recurrent neural networks with Gumbel-softmax output distributions in the task of generating sequences of discrete elements.
An Introduction to MM Algorithms for Machine Learning and Statistical
MM (majorization--minimization) algorithms are an increasingly popular tool for solving optimization problems in machine learning and statistical estimation. This article introduces the MM algorithm framework in general and via three popular example applications: Gaussian mixture regressions, multinomial logistic regressions, and support vector machines. Specific algorithms for the three examples are derived and numerical demonstrations are presented. Theoretical and practical aspects of MM algorithm design are discussed.
Use Azure Machine Learning with SQL Data Warehouse
Azure Machine Learning is a fully managed predictive analytics service that you can use to create predictive models against your data in SQL Data Warehouse, and then publish as ready-to-consume web services. You can learn the basics of predictive analytics and machine learning by reading Introduction to Machine Learning on Azure. You can then learn how to create, train, score and test a machine learning model using the Create experiment tutorial. We will read data from Product table in the AdventureWorksDW database. Start a new experiment by clicking NEW at the bottom of the Machine Learning Studio window, select EXPERIMENT, and then select Blank Experiment.
AI 'lawyer' correctly predicts outcomes of human rights trials
For the first time, artificial intelligence has been used to predict the outcomes of cases heard at a major European court. Researchers from the University of Sheffield, the University of Pennsylvania and University College London programmed the machine to analyse text from cases heard at the European Court of Human Rights (ECtHR) and predict the outcome of the judicial decision. During tests, the AI used a machine learning algorithm to make predictions with 79 per cent accuracy. "We don't see AI replacing judges or lawyers, but we think they'd find it useful for rapidly identifying patterns in cases that lead to certain outcomes," explained Dr Nikolaos Aletras, who led the study at UCL Computer Science. "It could also be a valuable tool for highlighting which cases are most likely to be violations of the European Convention on Human Rights." In developing the method, the team found judgements by the ECtHR correlate highly to non-legal facts rather than directly legal arguments, suggesting judges of the Court are, in the jargon of legal theory, 'realists' rather than'formalists'.
What if we are victims of an AI's singularity?
THIS is certainly the best book about the singularity. It features 26 intelligent scholars from 11 widely varying disciplines, all of them valiantly grappling with ghosts. Given that the subject matter is so highly speculative, so lofty, so indefinable, this tome is heavy going. Among its talents are nine philosophers and nine artificial intelligence researchers. These worthies mercilessly lay it on with their specialised jargon.
How Your Brain Decides Without You - Issue 42: Fakes
An autumn classic matching the unbeaten Tigers, with star tailback Dick Kazmaier--a gifted passer, runner, and punter who would capture a record number of votes to win the Heisman Trophy--against rival Dartmouth. Princeton prevailed over Big Green in the penalty-plagued game, but not without cost: Nearly a dozen players were injured, and Kazmaier himself sustained a broken nose and a concussion (yet still played a "token part"). It was a "rough game," The New York Times described, somewhat mildly, "that led to some recrimination from both camps." Each said the other played dirty. The game not only made the sports pages, it made the Journal of Abnormal and Social Psychology.
Annealing Gaussian into ReLU: a New Sampling Strategy for Leaky-ReLU RBM
Li, Chun-Liang, Ravanbakhsh, Siamak, Poczos, Barnabas
A BSTRACT Restricted Boltzmann Machine (RBM) is a bipartite graphical model that is used as the building block in energy-based deep generative models. Due to numerical stability and quantifiability of the likelihood, RBM is commonly used with Bernoulli units. Here, we consider an alternative member of exponential family RBM with leaky rectified linear units - called leaky RBM. We first study the joint and marginal distributions of leaky RBM under different leakiness, which provides us important insights by connecting the leaky RBM model and truncated Gaussian distributions. The connection leads us to a simple yet efficient method for sampling from this model, where the basic idea is to anneal the leakiness rather than the energy; - i.e., start from a fully Gaussian/Linear unit and gradually decrease the leakiness over iterations. This serves as an alternative to the annealing of the temperature parameter and enables numerical estimation of the likelihood that are more efficient and more accurate than the commonly used annealed importance sampling (AIS). We further demonstrate that the proposed sampling algorithm enjoys faster mixing property than contrastive divergence algorithm, which benefits the training without any additional computational cost. 1 I NTRODUCTION In this paper, we are interested in deep generative models. There is a family of directed deep generative models which can be trained by back-propagation (e.g., Kingma & Welling, 2013; Goodfellow et al., 2014). The other family is the deep energy-based models, including deep belief network (Hinton et al., 2006) and deep Boltzmann machine (Salakhutdinov & Hinton, 2009).
Recovery Guarantee of Non-negative Matrix Factorization via Alternating Updates
Li, Yuanzhi, Liang, Yingyu, Risteski, Andrej
Non-negative matrix factorization is a popular tool for decomposing data into feature and weight matrices under non-negativity constraints. It enjoys practical success but is poorly understood theoretically. This paper proposes an algorithm that alternates between decoding the weights and updating the features, and shows that assuming a generative model of the data, it provably recovers the ground-truth under fairly mild conditions. In particular, its only essential requirement on features is linear independence. Furthermore, the algorithm uses ReLU to exploit the non-negativity for decoding the weights, and thus can tolerate adversarial noise that can potentially be as large as the signal, and can tolerate unbiased noise much larger than the signal. The analysis relies on a carefully designed coupling between two potential functions, which we believe is of independent interest.
Understanding the 2016 US Presidential Election using ecological inference and distribution regression with census microdata
Flaxman, Seth, Sutherland, Dougal, Wang, Yu-Xiang, Teh, Yee Whye
The results of the 2016 US Presidential Election were, to put it mildly, a surprise. Pre-election polls and forecasts based on these polls pointed to a Clinton victory, a prediction shared by betting markets and pundits. In the aftermath of the vote, the main question asked is "why?" with answers ranging from the political to the economic to the social/cultural. In this article we attempt to provide a preliminary answer to a fundamental question: who voted for Trump, who voted for Clinton, and who voted for a third party or did not vote? By combining data from the United States census with the election results and recently proposed machine learning methods for ecological inference using regressions based on samples from a distribution, we provide local demographic estimates of voting and nonvoting. Unlike with exit polls, we are able to draw conclusions across the entire US and at a local level, about voters and non-voters, for interesting and novel combinations of predictor variables. It is our hope that this analysis will help inform the typical election post mortems, which are usually informed by incomplete information, due the following factors: - Vote counts will not be finalized in many precincts until days or in rare cases weeks after the election. Very close popular vote totals yield winner-take-all results, a fact of the US's electoral system but one that can lead to winner-take-all explanations.