Europe
How can we make the most of artificial intelligence? - Al Jazeera English
Machines and robots are increasingly shoving humans aside in the workplace. They build, cook and clean for us and increasingly work for us - so much so that Oxford University predicts almost half of US jobs will be done by machines and robots within the next 20 years. And then there is the next wave of technology: smarter artificial intelligence and augmented reality. It is a future that is leaving many people concerned and confused, while others see endless potential. What opportunities can artificial intelligence provide?
Tate Britain project uses AI to pair contemporary photos with paintings
Seated against a deep red backdrop, gazing intently at hand-held mirrors, two eunuchs in sparkling saris inspect their appearance before Raksha Bandhan celebrations in the red light district of Mumbai. The photograph from the Reuters news agency is an arresting contemporary scene, but a new Tate Britain project is aiming to inspire deeper reflections with images from its own collection of paintings. Launching on Friday, Recognition is the winner of 2016's IK prize – an annual award, this year supported by Microsoft, for a project that embraces digital technology to explore and showcase Tate's collection of British art. This year, the challenge was to do it with artificial intelligence. The team behind the winning project, from the Italy-based communication research centre Fabrica, say their inspiration came from an intriguing conundrum: how can you apply rational thinking to a subject like art? Recognition matches stunning photographs from the 24/7 news cycle with centuries-old artworks, and presents them online.
Google's DeepMind AI project apes human memory and programming skills
The mission of Google's DeepMind Technologies startup is to "solve intelligence." Now, researchers there have developed an artificial intelligence system that can mimic some of the brain's memory skills and even program like a human. The researchers developed a kind of neural network that can use external memory, allowing it to learn and perform tasks based on stored data. Neural networks are interconnected computational "neurons." While conventional neural networks have lacked readable and writeable memory, they have been used in machine learning and pattern-recognition applications such as computer vision and speech recognition.
The Future of Healthcare Is Arriving--8 Exciting Areas to Watch
The blending of home-based diagnostic platforms with medical care at home is arriving. The Tricorder XPRIZE competition is well underway, with several teams set to compete in the final stages. Leading contenders include CloudDx and Scanadu, a company started at our first Exponential Medicine program, have successfully leveraged crowdfunding to enable their clinical trials. Gale by 19Labs is a next generation "first aid kit meets home health center" (see the below video for a demo) exemplifying how integration of home diagnostics paired with menu-driven (and potentially AI-driven) assistance and optional telemedicine connectivity can provide increased access to home-based diagnosis, triage and management of minor bumps and scrapes and also more complex medical conditions. Interactive and engaging, from coaching on diet and nutrition to reminding you to take your medications or offering psychological support and follow up -- the chatbots are on their way.
Robust Discriminative Clustering with Sparse Regularizers
Flammarion, Nicolas, Palaniappan, Balamurugan, Bach, Francis
Clustering high-dimensional data often requires some form of dimensionality reduction, where clustered variables are separated from "noise-looking" variables. We cast this problem as finding a low-dimensional projection of the data which is well-clustered. This yields a one-dimensional projection in the simplest situation with two clusters, and extends naturally to a multi-label scenario for more than two clusters. In this paper, (a) we first show that this joint clustering and dimension reduction formulation is equivalent to previously proposed discriminative clustering frameworks, thus leading to convex relaxations of the problem, (b) we propose a novel sparse extension, which is still cast as a convex relaxation and allows estimation in higher dimensions, (c) we propose a natural extension for the multi-label scenario, (d) we provide a new theoretical analysis of the performance of these formulations with a simple probabilistic model, leading to scalings over the form $d=O(\sqrt{n})$ for the affine invariant case and $d=O(n)$ for the sparse case, where $n$ is the number of examples and $d$ the ambient dimension, and finally, (e) we propose an efficient iterative algorithm with running-time complexity proportional to $O(nd^2)$, improving on earlier algorithms which had quadratic complexity in the number of examples.
Visualizing and Understanding Sum-Product Networks
Vergari, Antonio, Di Mauro, Nicola, Esposito, Floriana
Sum-Product Networks (SPNs) are recently introduced deep tractable probabilistic models by which several kinds of inference queries can be answered exactly and in a tractable time. Up to now, they have been largely used as black box density estimators, assessed only by comparing their likelihood scores only. In this paper we explore and exploit the inner representations learned by SPNs. We do this with a threefold aim: first we want to get a better understanding of the inner workings of SPNs; secondly, we seek additional ways to evaluate one SPN model and compare it against other probabilistic models, providing diagnostic tools to practitioners; lastly, we want to empirically evaluate how good and meaningful the extracted representations are, as in a classic Representation Learning framework. In order to do so we revise their interpretation as deep neural networks and we propose to exploit several visualization techniques on their node activations and network outputs under different types of inference queries. To investigate these models as feature extractors, we plug some SPNs, learned in a greedy unsupervised fashion on image datasets, in supervised classification learning tasks. We extract several embedding types from node activations by filtering nodes by their type, by their associated feature abstraction level and by their scope. In a thorough empirical comparison we prove them to be competitive against those generated from popular feature extractors as Restricted Boltzmann Machines. Finally, we investigate embeddings generated from random probabilistic marginal queries as means to compare other tractable probabilistic models on a common ground, extending our experiments to Mixtures of Trees.
Wasserstein Discriminant Analysis
Flamary, Rémi, Cuturi, Marco, Courty, Nicolas, Rakotomamonjy, Alain
Wasserstein Discriminant Analysis (WDA) is a new supervised method that can improve classification of high-dimensional data by computing a suitable linear map onto a lower dimensional subspace. Following the blueprint of classical Linear Discriminant Analysis (LDA), WDA selects the projection matrix that maximizes the ratio of two quantities: the dispersion of projected points coming from different classes, divided by the dispersion of projected points coming from the same class. To quantify dispersion, WDA uses regularized Wasserstein distances, rather than cross-variance measures which have been usually considered, notably in LDA. Thanks to the the underlying principles of optimal transport, WDA is able to capture both global (at distribution scale) and local (at samples scale) interactions between classes. Regularized Wasserstein distances can be computed using the Sinkhorn matrix scaling algorithm; We show that the optimization of WDA can be tackled using automatic differentiation of Sinkhorn iterations. Numerical experiments show promising results both in terms of prediction and visualization on toy examples and real life datasets such as MNIST and on deep features obtained from a subset of the Caltech dataset.
Discovering Patterns in Time-Varying Graphs: A Triclustering Approach
Guigourès, Romain, Boullé, Marc, Rossi, Fabrice
This paper introduces a novel technique to track structures in time varying graphs. The method uses a maximum a posteriori approach for adjusting a three-dimensional co-clustering of the source vertices, the destination vertices and the time, to the data under study, in a way that does not require any hyper-parameter tuning. The three dimensions are simultaneously segmented in order to build clusters of source vertices, destination vertices and time segments where the edge distributions across clusters of vertices follow the same evolution over the time segments. The main novelty of this approach lies in that the time segments are directly inferred from the evolution of the edge distribution between the vertices, thus not requiring the user to make any a priori quantization. Experiments conducted on artificial data illustrate the good behavior of the technique, and a study of a real-life data set shows the potential of the proposed approach for exploratory data analysis.
Paris Machine Learning Newsletter, Summer 2016
We've had more than 150 speakers in the past three seasons. Two of them made the news this summer: Danny Bickson (E9 Season 1) one of the co-founders of Graphlab then Dato then Turi and Arjun Bansal from Nervana systems (E12 Season 3). Turi just got acquired by Apple for 300M, and Nervana got acquired for 350M by Intel. In a different direction, at the last meetup, Raymond Francis explained to us what got picked by the LA Times a month later, Curiosity now uses Machine Learning on Mars. This news is exciting on two levels: First, robots can now explore the universe better and second, it definitely brings some perspective when we talk about the dichotomy between exploration and exploitation in our discussions.
TragiComedy hour: P-values vs posterior probabilities vs diagnostic error rates
The consequences of recent criticisms of statistical tests have breathed brand new life into some very old howlers, many of which have been discussed on this blog. What is not funny, though, is how standard notions such as frequentist error probabilities are being redefined in the process, and how we now have arguments built on equivocations. In fact, there are official guidebooks for the statistically perplexed giving inconsistent definitions to the same term (See for just 1 of many examples this post). How much more perplexed will that leave us! Since it's near the 5-year anniversary of this blog, tonight let's listen in to a new comedy hour mixing one from 3 years ago with some add-ons*.