Goto

Collaborating Authors

 Europe


Markov Properties for Graphical Models with Cycles and Latent Variables

arXiv.org Machine Learning

We investigate probabilistic graphical models that allow for both cycles and latent variables. For this we introduce directed graphs with hyperedges (HEDGes), generalizing and combining both marginalized directed acyclic graphs (mDAGs) that can model latent (dependent) variables, and directed mixed graphs (DMGs) that can model cycles. We define and analyse several different Markov properties that relate the graphical structure of a HEDG with a probability distribution on a corresponding product space over the set of nodes, for example factorization properties, structural equations properties, ordered/local/global Markov properties, and marginal versions of these. The various Markov properties for HEDGes are in general not equivalent to each other when cycles or hyperedges are present, in contrast with the simpler case of directed acyclic graphical (DAG) models (also known as Bayesian networks). We show how the Markov properties for HEDGes - and thus the corresponding graphical Markov models - are logically related to each other.


Learning Depth-Three Neural Networks in Polynomial Time

arXiv.org Machine Learning

We give a polynomial-time algorithm for learning neural networks with one hidden layer of sigmoids feeding into any smooth, monotone activation function (e.g., sigmoid or ReLU). We make no assumptions on the structure of the network, and the algorithm succeeds with respect to {\em any} distribution on the unit ball in $n$ dimensions (hidden weight vectors also have unit norm). This is the first assumption-free, provably efficient algorithm for learning neural networks with more than one hidden layer. Our algorithm-- {\em Alphatron}-- is a simple, iterative update rule that combines isotonic regression with kernel methods. It outputs a hypothesis that yields efficient oracle access to interpretable features. It also suggests a new approach to Boolean function learning via smooth relaxations of hard thresholds, sidestepping traditional hardness results from computational learning theory. Along these lines, we give improved results for a number of longstanding problems related to Boolean concept learning, unifying a variety of different techniques. For example, we give the first polynomial-time algorithm for learning intersections of halfspaces with a margin (distribution-free) and the first generalization of DNF learning to the setting of probabilistic concepts (queries; uniform distribution). Finally, we give the first provably correct algorithms for common schemes in multiple-instance learning.


Distributionally Ambiguous Optimization Techniques in Batch Bayesian Optimization

arXiv.org Machine Learning

We propose a novel, theoretically-grounded, acquisition function for batch Bayesian optimization informed by insights from distributionally ambiguous optimization. Our acquisition function is a lower bound on the well-known Expected Improvement function -- which requires a multi-dimensional Gaussian Expectation over a piecewise affine function -- and is computed by evaluating instead the best-case expectation over all probability distributions consistent with the same mean and variance as the original Gaussian distribution. Unlike alternative approaches including Expected Improvement, our proposed acquisition function avoids multi-dimensional integrations entirely, and can be computed exactly as the solution of a convex optimization problem in the form of a tractable semidefinite program (SDP). Moreover, we prove that the solution of this SDP also yields exact numerical derivatives, which enable efficient optimization of the acquisition function. Finally, it efficiently handles marginalized posteriors with respect to the Gaussian Process' hyperparameters. We demonstrate superior performance to heuristic alternatives and approximations of the intractable expected improvement, justifying this performance difference based on simple examples that break the assumptions of state-of-the-art methods.


Learning how to explain neural networks: PatternNet and PatternAttribution

arXiv.org Machine Learning

DeConvNet, Guided BackProp, LRP, were invented to better understand deep neural networks. We show that these methods do not produce the theoretically correct explanation for a linear model. Yet they are used on multi-layer networks with millions of parameters. This is a cause for concern since linear models are simple neural networks. We argue that explanation methods for neural nets should work reliably in the limit of simplicity, the linear models. Based on our analysis of linear models we propose a generalization that yields two explanation techniques (PatternNet and PatternAttribution) that are theoretically sound for linear models and produce improved explanations for deep networks.


A Mathematical Theory of Deep Convolutional Neural Networks for Feature Extraction

arXiv.org Artificial Intelligence

Deep convolutional neural networks have led to breakthrough results in numerous practical machine learning tasks such as classification of images in the ImageNet data set, control-policy-learning to play Atari games or the board game Go, and image captioning. Many of these applications first perform feature extraction and then feed the results thereof into a trainable classifier. The mathematical analysis of deep convolutional neural networks for feature extraction was initiated by Mallat, 2012. Specifically, Mallat considered so-called scattering networks based on a wavelet transform followed by the modulus non-linearity in each network layer, and proved translation invariance (asymptotically in the wavelet scale parameter) and deformation stability of the corresponding feature extractor. This paper complements Mallat's results by developing a theory that encompasses general convolutional transforms, or in more technical parlance, general semi-discrete frames (including Weyl-Heisenberg filters, curvelets, shearlets, ridgelets, wavelets, and learned filters), general Lipschitz-continuous non-linearities (e.g., rectified linear units, shifted logistic sigmoids, hyperbolic tangents, and modulus functions), and general Lipschitz-continuous pooling operators emulating, e.g., sub-sampling and averaging. In addition, all of these elements can be different in different network layers. For the resulting feature extractor we prove a translation invariance result of vertical nature in the sense of the features becoming progressively more translation-invariant with increasing network depth, and we establish deformation sensitivity bounds that apply to signal classes such as, e.g., band-limited functions, cartoon functions, and Lipschitz functions.


ISS Astronauts Operating Remote Robots Show Future of Planetary Exploration

IEEE Spectrum Robotics

In late August, an astronaut on board the International Space Station remotely operated a humanoid robot to inspect and repair a solar farm on Mars--or at least a simulated Mars environment, which in this case is a room with rust-colored floors, walls, and curtains at the German Aerospace Center, or DLR, in Oberpfaffenhofen, near Munich. European Space Agency astronaut Paolo Nespoli commanded the humanoid, called Rollin' Justin, as the robot performed a series of navigation, maintenance, and repair tasks. Instead of relying on direct teleoperation, Nespoli used a tablet computer to issue high-level commands to the robot. In one task, he used the tablet to position the robot and have it take pictures from different angles. Another command instructed Justin to grasp a cable and connect it to a data port.


Robotic underwater miners can go where humans can't

New Scientist

The scene around the flooded Whitehill Yeo pit in Devon, UK, resembles a lunar landscape. Until it was abandoned just a few years ago, an endless stream of diesel trucks carried china clay out of the mine seven days a week. But don't be fooled by the silence: this is very much an active site. It's just that all the excavation is happening deep beneath the placid waters. This is a test bed, the first, for a new type of mining by underwater robots.


Can Crowdsourcing Teach AI to Do the Right Thing?

#artificialintelligence

As we cede more and more control to artificial intelligence, it's inevitable that those machines will need to make choices based, hopefully, on human morality. But where are AI's ethics going to come from? Could they be crowdsourced, essentially voted on by everyone? Alphabet's DeepMind division now has a unit working on AI ethics, and in June 2017, Germany became the first nation to officially begin to address the question, with a report issued by its Ethics Commission on Automated and Connected Driving. For anyone worried about machines taking over -- to quote Stephen Hawking, "The development of artificial intelligence could spell the end of the human race."


Amazon to open visually focused AI research hub in Germany

#artificialintelligence

Ecommerce giant Amazon has announced a new research center in Germany focused on developing AI to improve the customer experience -- especially in visual systems. Amazon said research conducted at the hub will also aim to benefit users of Amazon Web Services and its voice driven AI assistant tech, Alexa. The center will be based in Tübingen, near the Max Planck Institute for Intelligent Systems' campus, and will be staffed with more than 100 machine learning engineers. The new 100 "highly qualified" jobs will be created over the next five years, it said today. The site is Amazon's fourth Research Center in Germany -- after Berlin, Dresden and Aachen.


Driverless pods could replace night buses in Cambridge

Daily Mail - Science & tech

Driverless robotic pods are being tested on Cambridge busways in the hope they could solve Britain's traffic problems. The futuristic vehicles, which can carry four people at a time, could pave the way for an evening public transport service in Cambridge. But anyone in a hurry may need to brace themselves – the futuristic vehicles have a top speed of just 15mph (24km/h). Driverless robotic pods called Podzeros are being tested on busways in Cambridge in the hope they could solve Britain's traffic problems The futuristic vehicles, which can carry four people at a time, are being trialled as an alternative method of public transport. It's hoped they would ease congestion by offering an automated service to make around 100 journeys a day - away from pedestrians and cyclists.