Europe
Chris Bishop elected as a Fellow of the Royal Society - The AI Blog
Chris Bishop, a world-renowned expert in artificial intelligence and machine learning, was elected a Fellow of the Royal Society, the oldest scientific academy in continuous existence, on Friday. It's a pinnacle of my career," said Bishop, distinguished scientist and director of Microsoft Research Cambridge, the European arm of Microsoft's research organization. Bishop is among 50 other scientists from across the United Kingdom and Commonwealth and 10 Foreign Members elected to the Royal Society, including pioneers in understanding the chemical origins of life, and discovering how humans operate on a 24-hour cycle. "Science is a great triumph of human achievement and has contributed hugely to the prosperity and health of our world. In the coming decades it will play an increasingly crucial role in tackling the great challenges of our time including food, energy, health and the environment. The new Fellows of the Royal Society have already contributed much to science and it gives me great pleasure to welcome them into our ranks," said Venki Ramakrishnan, President of the Royal Society, in announcing the news earlier today.
These Seven Countries Are In A Race To Rule The World With AI
In a recent speech, Russian president Vladimir Putin made an incredibly prescient statement: "Artificial intelligence is the future, not only for Russia, but for all of humankind." He went on to highlight both the risks and rewards of AI and concluded by declaring that whatever country comes to dominate this technology will be the "ruler of the world." As someone who closely monitors global events and studies emerging technologies, I think Putin's lofty rhetoric is entirely appropriate. Funding for global AI startups has grown at a 60% compound annual growth rate since 2010. More significantly, the international community is actively discussing the influence AI will exert over both global cooperations and national strength.
Google brings voice calling to Home speakers in the UK
That means any of your flatmates or family members can say "call mum" and get the right person. Google's Assistant can also find the number for "millions" of businesses across Britain, so phoning your local pharmacy, mechanic or pizzeria should be a breeze. When you call someone for the first time, your Home speaker will show up as an unknown or private number. It's a pain -- especially if you're trying to call grandma -- but you can set up caller ID so your number appears for every subsequent call. To celebrate Mother's Day (in the UK, anyway) Google is dropping the price of its coral-colored Home Mini speaker by £10 to £39 until March 12th.
Self-driving cars attacked by angry San Francisco residents
Technology and automotive companies touting self-driving cars as the future of transportation may have some work to convince San Franciscans, who keep attacking the vehicles. A third of traffic collisions involving autonomous vehicles in 2018 so far featured humans physically confronting the cars, according to data released by California. In one case, a taxi driver exited his cab and slapped the front passenger window of a General Motors Cruise parked behind him. No one was hurt, though the car sustained a scratch. In another case, a pedestrian hurtled across an intersection despite a "do not walk" sign, shouting as he went, and rammed his body into a different Cruise's rear bumper.
Neural-Network Quantum States, String-Bond States, and Chiral Topological States
Glasser, Ivan, Pancotti, Nicola, August, Moritz, Rodriguez, Ivan D., Cirac, J. Ignacio
Neural-Network Quantum States have been recently introduced as an Ansatz for describing the wave function of quantum many-body systems. We show that there are strong connections between Neural-Network Quantum States in the form of Restricted Boltzmann Machines and some classes of Tensor-Network states in arbitrary dimensions. In particular we demonstrate that short-range Restricted Boltzmann Machines are Entangled Plaquette States, while fully connected Restricted Boltzmann Machines are String-Bond States with a nonlocal geometry and low bond dimension. These results shed light on the underlying architecture of Restricted Boltzmann Machines and their efficiency at representing many-body quantum states. String-Bond States also provide a generic way of enhancing the power of Neural-Network Quantum States and a natural generalization to systems with larger local Hilbert space. We compare the advantages and drawbacks of these different classes of states and present a method to combine them together. This allows us to benefit from both the entanglement structure of Tensor Networks and the efficiency of Neural-Network Quantum States into a single Ansatz capable of targeting the wave function of strongly correlated systems. While it remains a challenge to describe states with chiral topological order using traditional Tensor Networks, we show that Neural-Network Quantum States and their String-Bond States extension can describe a lattice Fractional Quantum Hall state exactly. In addition, we provide numerical evidence that Neural-Network Quantum States can approximate a chiral spin liquid with better accuracy than Entangled Plaquette States and local String-Bond States. Our results demonstrate the efficiency of neural networks to describe complex quantum wave functions and pave the way towards the use of String-Bond States as a tool in more traditional machine-learning applications.
A strong converse bound for multiple hypothesis testing, with applications to high-dimensional estimation
Venkataramanan, Ramji, Johnson, Oliver
In statistical language we seek to give a lower bound on the performance of any estimator over a class of problems (often called the minimax risk over the class). In the language of information theory, we speak of converse results, which give performance bounds for all communication schemes over a noisy channel. In the statistics literature, one standard approach to proving converse results is via Fano's inequality (see [1, Theorem 2.11.1]). However, recent information-theoretic literature has shown how to obtain sharper converse bounds. The resulting improvements can be significant at finite sample size, and give bounds that are close to optimal, as illustrated in the work of Polyanskiy, Poor and Verdú [2]. The present paper shows how the method of [2], although developed for channel coding problems, gives stronger risk lower bounds for high-dimensional estimation problems, compared to the standard Fano approach. We first describe the general setup, following the treatment and notation of [3, Chapter 2].
Distributed Computation of Wasserstein Barycenters over Networks
Uribe, César A., Dvinskikh, Darina, Dvurechensky, Pavel, Gasnikov, Alexander, Nedić, Angelia
Optimal Transport distances (also known as earth mover's distances or Wasserstein distances) design an optimal plan to move "mass" from one probability distribution to another. This problem can be traced back to the early work of Monge [1] and Kantorovich [2] and has been of constant interest for allowing natural formulations to the problems of comparing, interpolating, and measuring distances of functions [3]. On the other hand, computational optimal transport has gain popularity for its applications in learning theory [4], computer vision [5], computer graphics [6], statistical inference [7], information fusion [8]; and its relative complexity advantages with respect to classical methods [9]. Particularly, large-scale optimal transport has been of recent interest for the latest applications where large quantities of data are available and efficient algorithms are required [10, 11, 12]. Comprehensive accounts of the optimal transport problem and its computational aspects can be found in [13, 14, 15, 3]. One of the common uses of the Wasserstein distance is the aggregation of distributions by considering their barycenter [16], which itself is another distribution [17]. Wasserstein Barycenters has been shown superior to traditional Euclidean-based methods in a range of application such as image processing [16], economics and finance [18] and condensed matter physics [19]. Figure 1 shows a sample of 100 images of the digit 7 from the MNIST dataset [20] and their respective Euclidean mean and Wasserstein mean.
Stochastic Block Models with Multiple Continuous Attributes
Stanley, Natalie, Bonacci, Thomas, Kwitt, Roland, Niethammer, Marc, Mucha, Peter J.
Abstract--The stochastic block model (SBM) is a probabilistic model for community structure in networks. Typically, only the adjacency matrix is used to perform SBM parameter inference. In this paper, we consider circumstances in which nodes have an associated vector of continuous attributes that are also used to learn the node-to-community assignments and corresponding SBM parameters. While this assumption is not realistic for every application, our model assumes that the attributes associated with the nodes in a network's community can be described by a common multivariate Gaussian model. In this augmented, attributed SBM, the objective is to simultaneously learn the SBM connectivity probabilities with the multivariate Gaussian parameters describing each community. While there are recent examples in the literature that combine connectivity and attribute information to inform community detection, our model is the first augmented stochastic block model to handle multiple continuous attributes. This provides the flexibility in biological data to, for example, augment connectivity information with continuous measurements from multiple experimental modalities. Because the lack of labeled network data often makes community detection results difficult to validate, we highlight the usefulness of our model for two network prediction tasks: link prediction and collaborative filtering. As a result of fitting this attributed stochastic block model, one can predict the attribute vector or connectivity patterns for a new node in the event of the complementary source of information (connectivity or attributes, respectively). We also highlight two biological examples where the attributed stochastic block model provides satisfactory performance in the link prediction and collaborative filtering tasks. In various applications, each node in a network is equipped with additional information (or particular attributes) that was not implicitly taken into account in the construction of the network.
A Walk with SGD
Xing, Chen, Arpit, Devansh, Tsirigotis, Christos, Bengio, Yoshua
Exploring why stochastic gradient descent (SGD) based optimization methods train deep neural networks (DNNs) that generalize well has become an active area of research. Towards this end, we empirically study the dynamics of SGD when training over-parametrized DNNs. Specifically we study the DNN loss surface along the trajectory of SGD by interpolating the loss surface between parameters from consecutive \textit{iterations} and tracking various metrics during training. We find that the loss interpolation between parameters before and after a training update is roughly convex with a minimum (\textit{valley floor}) in between for most of the training. Based on this and other metrics, we deduce that during most of the training, SGD explores regions in a valley by bouncing off valley walls at a height above the valley floor. This 'bouncing off walls at a height' mechanism helps SGD traverse larger distance for small batch sizes and large learning rates which we find play qualitatively different roles in the dynamics. While a large learning rate maintains a large height from the valley floor, a small batch size injects noise facilitating exploration. We find this mechanism is crucial for generalization because the valley floor has barriers and this exploration above the valley floor allows SGD to quickly travel far away from the initialization point (without being affected by barriers) and find flatter regions, corresponding to better generalization.
Towards better understanding of gradient-based attribution methods for Deep Neural Networks
Ancona, Marco, Ceolini, Enea, Öztireli, Cengiz, Gross, Markus
Understanding the flow of information in Deep Neural Networks (DNNs) is a challenging problem that has gain increasing attention over the last few years. While several methods have been proposed to explain network predictions, there have been only a few attempts to compare them from a theoretical perspective. What is more, no exhaustive empirical comparison has been performed in the past. In this work, we analyze four gradient-based attribution methods and formally prove conditions of equivalence and approximation between them. By reformulating two of these methods, we construct a unified framework which enables a direct comparison, as well as an easier implementation. Finally, we propose a novel evaluation metric, called Sensitivity-n and test the gradient-based attribution methods alongside with a simple perturbation-based attribution method on several datasets in the domains of image and text classification, using various network architectures.