Deep Learning
Before artificial intelligence we need to understand human intelligence
In the global race to build artificial intelligence, it was a missed opportunity. Jeff Hawkins, a Silicon Valley veteran who spent the last decade exploring the mysteries of the human brain, arranged a meeting with DeepMind, the world's leading AI lab. Scientists at DeepMind, which is owned by Google's parent company, Alphabet, want to build machines that can do anything the brain can do. Hawkins runs a little company with one goal: figure out how the brain works and then reverse engineer it. The meeting, which had been set for April at DeepMind's offices in London, never happened.
AI In Medicine: Rise Of The Machines
Could a robot do my job as a radiologist? If you asked me 10 years ago, I would have said, "No way!" But if you ask me today, my answer would be more hesitant, "Not yet -- but perhaps someday soon." In particular, new "deep learning" artificial intelligence (AI) algorithms are showing promise in performing medical work which until recently was thought only capable of being done by human physicians. For example, deep learning algorithms have been able to diagnose the presence or absence of tuberculosis (TB) in chest x-ray images with astonishing accuracy.
Deep Learning Summit Coming to Toronto in October
The Deep Learning Summit is returning to Toronto from October 25 โ 26, 2018 and will cover the latest advancements in deep learning technology. Global leaders in the field will address how industry leaders and start-ups are applying deep learning techniques across industry and society. The first ever AI for Government Summit, another event stream, will provide a unique opportunity to interact with government bodies, policymakers, strategists and directors of innovation to explore the use of machine learning to increase efficiency, reduce costs and meet the high demands of the public sector. What's more, the Canadian Government have committed over $125 million to AI developments. Headline partners include Accenture, Qualcomm, Graphcore AI and CBC/Radio Canada who will all be sharing their expertise in the field, participating in workshops, discussions, presentations, demonstrations and exhibitions.
Restricted Boltzmann Machines -- Simplified โ Towards Data Science
In this post, I will try to shed some light on the intuition about Restricted Boltzmann Machines and the way they work. This is supposed to be a simple explanation with a little bit of mathematics without going too deep into each concept or equation. So let's start with the origin of RBMs and delve deeper as we move forward. Boltzmann machines are stochastic and generative neural networks capable of learning internal representations and are able to represent and (given sufficient time) solve difficult combinatoric problems. They are named after the Boltzmann distribution (also known as Gibbs Distribution) which is an integral part of Statistical Mechanics and helps us to understand the impact of parameters like Entropy and Temperature on the Quantum States in Thermodynamics.
How AI Could Save a Submarine from Attack
The underwater ocean world is an ecosystem with lots of different sounds. So naval forces have traditionally relied on so-called "golden ears," or musicians and other individuals with particularly sharp hearing, to detect the specific signals coming from an enemy submarine. But given the overload of data today, distinguishing between false alarms and actual dangers has become more difficult. That's why "Thales is working on "Deep Learning" algorithms capable of recognizing the particular "song" of a submarine, much as the "Shazam" app helps you identify a song you hear on the radio", says Dominique Thubert, Thales Underwater Systems, which is specialized in sonar systems for submarines, surface warships, and aircraft. These algorithms, attached to submarines, surface ship or drones, will help naval forces sort through and classify information in order to detect attacks early on.
Diagnosing Lung Disease Using Deep Learning - Intel AI
Research Using CheXNet at Stanford: CheXNet is a deep learning Convolutional Neural Network (CNN) model developed at Stanford University to identify thoracic pathologies from the NIH ChestXray14 dataset. CheXNet is a 121-layer CNN that uses chest X-Ray images to predict the output probabilities of a pathology. It correctly detects pneumonia by localizing the areas in the image that are most indicative of the pathology. Stanford researchers have been able to train the ChestX-Ray14 dataset using a pre-trained model of CheXNet-121 with the ImageNet2012-1K dataset. The NIH dataset consists of over one hundred thousand frontal chest X-ray images from over 30,000 unique patients that have been annotated with up to 14 thoracic diseases including pneumonia and emphysema.
MMLSpark: Unifying Machine Learning Ecosystems at Massive Scales
Hamilton, Mark, Raghunathan, Sudarshan, Matiach, Ilya, Schonhoffer, Andrew, Raman, Anand, Barzilay, Eli, Thigpen, Minsoo, Rajendran, Karthik, Mahajan, Janhavi Suresh, Cochrane, Courtney, Eswaran, Abhiram, Green, Ari
We introduce Microsoft Machine Learning for Apache Spark (MMLSpark), an ecosystem of enhancements that expand the Apache Spark distributed computing library to tackle problems in Deep Learning, Micro-Service Orchestration, Gradient Boosting, Model Interpretability, and other areas of modern computation. Furthermore, we present a novel system called Spark Serving that allows users to run any Apache Spark program as a distributed, sub-millisecond latency web service backed by their existing Spark Cluster. All MMLSpark contributions have the same API to enable simple composition across frameworks and usage across batch, streaming, and RESTful web serving scenarios on static, elastic, or serverless clusters. We showcase MMLSpark by creating a method for deep object detection capable of learning without human labeled data and demonstrate its effectiveness for Snow Leopard conservation.
CNN inference acceleration using dictionary of centroids
Babin, D., Mazurenko, I., Parkhomenko, D., Voloshko, A.
It is well known that multiplication operations in convolutional layers of common CNNs consume a lot of time during inference stage. In this article we present a flexible method to decrease both computational complexity of convolutional layers in inference as well as amount of space to store them. The method is based on centroid filter quantization and outperforms approaches based on tensor decomposition by a large margin. We performed comparative analysis of the proposed method and series of CP tensor decomposition on ImageNet benchmark and found that our method provide almost 2.9 times better computational gain. Despite the simplicity of our method it cannot be applied directly in inference stage in modern frameworks, but could be useful for cases calculation flow could be changed, e.g. for CNN-chip designers.
A Modern Take on the Bias-Variance Tradeoff in Neural Networks
Neal, Brady, Mittal, Sarthak, Baratin, Aristide, Tantia, Vinayak, Scicluna, Matthew, Lacoste-Julien, Simon, Mitliagkas, Ioannis
We revisit the bias-variance tradeoff for neural networks in light of modern empirical findings. The traditional bias-variance tradeoff in machine learning suggests that as model complexity grows, variance increases. Classical bounds in statistical learning theory point to the number of parameters in a model as a measure of model complexity, which means the tradeoff would indicate that variance increases with the size of neural networks. However, we empirically find that variance due to training set sampling is roughly constant (with both width and depth) in practice. Variance caused by the non-convexity of the loss landscape is different. We find that it decreases with width and increases with depth, in our setting. We provide theoretical analysis, in a simplified setting inspired by linear models, that is consistent with our empirical findings for width. We view bias-variance as a useful lens to study generalization through and encourage further theoretical explanation from this perspective. The traditional view in machine learning is that increasingly complex models achieve lower bias at the expense of higher variance. This balance between underfitting (high bias) and overfitting (high variance) is commonly known as the bias-variance tradeoff (Figure 1). In their landmark work that initially highlighted this bias-variance dilemma in machine learning, Geman et al. (1992) suggest that larger neural networks suffer from higher variance.
Nonparametric Bayesian Lomax delegate racing for survival analysis with competing risks
We propose Lomax delegate racing (LDR) to explicitly model the mechanism of survival under competing risks and to interpret how the covariates accelerate or decelerate the time to event. LDR explains non-monotonic covariate effects by racing a potentially infinite number of sub-risks, and consequently relaxes the ubiquitous proportional-hazards assumption which may be too restrictive. Moreover, LDR is naturally able to model not only censoring, but also missing event times or event types. For inference, we develop a Gibbs sampler under data augmentation for moderately sized data, along with a stochastic gradient descent maximum a posteriori inference algorithm for big data applications. Illustrative experiments are provided on both synthetic and real datasets, and comparison with various benchmark algorithms for survival analysis with competing risks demonstrates distinguished performance of LDR.