Deep Learning
DeepMind: First major AI patent filings revealed
DeepMind, a leading artificial intelligence (AI) research company, has filed a series of international patent applications, which have now been published for the first time. The applications relate to a number of the fundamental aspects of modern day machine learning, and are therefore of potential significance to anyone operating in the commercial AI sector. Background to DeepMind DeepMind is a London based artificial intelligence (AI) research company, widely recognized as being at the forefront of the field. DeepMind was founded in 2010 and acquired by Google in 2014 for £400m. In 2017, DeepMind famously developed AI capable of defeating a world-champion at Go (Silver et al.
Artificial Intelligence Better Than Doctors At Diagnosing Skin Cancer
Skin cancer was found to be diagnosed more accurately by artificial intelligence than experienced dermatologists in a new international study. Researchers tested a form of machine learning known as a deep learning convolutional neural network (CNN) to reach this conclusion. The study titled "Artificial intelligence for melanoma diagnosis: How can we deliver on the promise?" was published in the cancer journal Annals of Oncology on May 28. Malignant melanoma accounts for 1 percent of all skin cancers but causes a majority of skin cancer-related deaths. The American Cancer Society estimates 9,320 people will die from melanoma in 2018 while 91,270 new cases will be diagnosed.
Demystifying Deep Learning - Back to Basics Vinod Sharma's Blog
This is part 1 of 2 parts story on DeepLearning & basic terms which revolve (may evolve around as well) around it. Deep Learning is a very young field, where theories aren't strongly established and views quickly changes almost on daily basis. Deep Learning is at the cutting edge technology break through. This depicts what machines can do (still very new and at basic level), and developers and business leaders absolutely need to understand what it is and how it works. "I think people needs to understand that deep learning is making a lot of things, behind the scenes, much better" – Sir Geoffrey Hinton With lots of noise I can say "Deep learning is undeniably mind-blowing" and "deep learning can be used with too much of ease to predict the unpredictable".
ACCESS GRANTED – Tomorrow's Business Ethics
Industry 4.0 with the Internet of Things and Artificial Intelligence (including Machine and Deep Learning) started to interrupt today's reality, including the way we conduct business. It was a pleasure to take a look into the near future and predict the coming changes. The book was not the destination, but the start of the journey, as I had the honor to discuss the different topics with acknowledged experts and receive priceless feedback. Artificial Intelligence, chat-bots & robots, 3D printing, micro-learnings, virtual & augmented reality, self-driving cars and all other autonomous software & machines will be a part of tomorrow's business. We have to start thinking about the consequences. A chance and challenge for management, where the Ethics&Compliance-department can position itself as a key-player and include AI inside its responsibilities.
MIT Creates Norman, the World's First Psychopath Artificial Intelligence System - TechEBlog
Massachusetts Institute of Technology researchers have managed to train an artificial intelligence algorithm, called "Norman", to become a psychopath by exposing it to gruesome or violent images posted on Reddit. Named after the Anthony Perkins character in Alfred Hitchcock's 1960 classic "Psycho", this AI was trained to perform image captioning, a popular deep learning method of generating a textual description of an image, but the twist was only exposing it to gruesome and violent images from a subreddit dedicated to documenting and observing the disturbing reality of death. Rorsach inkblots were then used to compare Norman to other AI which hadn't been exposed to the same gruesome images. "Norman is born from the fact that the data that is used to teach a machine learning algorithm can significantly influence its behavior. So when people talk about AI algorithms being biased and unfair, the culprit is often not the algorithm itself, but the biased data that was fed to it. The same method can see very different things in an image, even sick things, if trained on the wrong (or, the right!) data set. Norman suffered from extended exposure to the darkest corners of Reddit, and represents a case study on the dangers of Artificial Intelligence gone wrong when biased data is used in machine learning algorithms," said MIT scientists.
May 23 Webinar: How to Consume Your Data for AI - DATAVERSITY
Smarter businesses apply AI to learn and continuously evolve the way they work. To extract full value from AI, companies need data strategy that gives them access to all their data – no matter where it lives – in an environment that easily scales and applies the latest discovery technology including advanced analytics, visualization and AI. Learn how IBM Watson and Data provides all the tools companies need to embed AI, machine learning and deep learning in their business, while enabling professionals to gain the most from their data to drive smarter business and lead industry-changing transformations.
Auto-Meta: Automated Gradient Based Meta Learner Search
Kim, Jaehong, Choi, Youngduck, Cha, Moonsu, Lee, Jung Kwon, Lee, Sangyeul, Kim, Sungwan, Choi, Yongseok, Kim, Jiwon
Fully automating machine learning pipeline is one of the outstanding challenges of general artificial intelligence, as practical machine learning often requires costly human driven process, such as hyper-parameter tuning, algorithmic selection, and model selection. In this work, we consider the problem of executing automated, yet scalable search for finding optimal gradient based meta-learners in practice. As a solution, we apply progressive neural architecture search to proto-architectures by appealing to the model agnostic nature of general gradient based meta learners. In the presence of recent universality result of Finn \textit{et al.}\cite{finn:universality_maml:DBLP:/journals/corr/abs-1710-11622}, our search is a priori motivated in that neural network architecture search dynamics---automated or not---may be quite different from that of the classical setting with the same target tasks, due to the presence of the gradient update operator. A posteriori, our search algorithm, given appropriately designed search spaces, finds gradient based meta learners with non-intuitive proto-architectures that are narrowly deep, unlike the inception-like structures previously observed in the resulting architectures of traditional NAS algorithms. Along with these notable findings, the searched gradient based meta-learner achieves state-of-the-art results on the few shot classification problem on Mini-ImageNet with $76.29\%$ accuracy, which is an $13.18\%$ improvement over results reported in the original MAML paper. To our best knowledge, this work is the first successful AutoML implementation in the context of meta learning.
Gear Training: A new way to implement high-performance model-parallel training
Dong, Hao, Li, Shuai, Xu, Dongchang, Ren, Yi, Zhang, Di
The training of Deep Neural Networks usually needs tremendous computing resources. Therefore many deep models are trained in large cluster instead of single machine or GPU. Though major researchs at present try to run whole model on all machines by using asynchronous asynchronous stochastic gradient descent (ASGD)[9], we present a new approach to train deep model parallely - split the model and then seperately train different parts of it in different speed.
Chaining Mutual Information and Tightening Generalization Bounds
Asadi, Amir R., Abbe, Emmanuel, Verdú, Sergio
Bounding the generalization error of learning algorithms has a long history, that yet falls short in explaining various generalization successes including those of deep learning. Two important difficulties are (i) exploiting the dependencies between the hypotheses, (ii) exploiting the dependence between the algorithm's input and output. Progress on the first point was made with the chaining method, originating from the work of Kolmogorov and used in the VC-dimension bound. More recently, progress on the second point was made with the mutual information method by Russo and Zou '15. Yet, these two methods are currently disjoint. In this paper, we introduce a technique to combine chaining and mutual information methods, to obtain a generalization bound that is both algorithm-dependent and that exploits the dependencies between the hypotheses. We provide an example in which our bound significantly outperforms both the chaining and the mutual information bounds. As a corollary, we tighten Dudley inequality under the knowledge that a learning algorithm chooses its output from a small subset of hypotheses with high probability; an assumption motivated by the performance of SGD discussed in Zhang et al. '17.
A Note about: Local Explanation Methods for Deep Neural Networks lack Sensitivity to Parameter Values
Sundararajan, Mukund, Taly, Ankur
Local explanation methods, also known as attribution methods, attribute a deep network's prediction to its input (cf. Baehrens et al. (2010)). We respond to the claim from Adebayo et al. (2018) that local explanation methods lack sensitivity, i.e., DNNs with randomly-initialized weights produce explanations that are both visually and quantitatively similar to those produced by DNNs with learned weights. Further investigation reveals that their findings are due to two choices in their analysis: (a) ignoring the signs of the attributions; and (b) for integrated gradients (IG), including pixels in their analysis that have zero attributions by choice of the baseline (an auxiliary input relative to which the attributions are computed). When both factors are accounted for, IG attributions for a random network and the actual network are uncorrelated. Our investigation also sheds light on how these issues affect visualizations, although we note that more work is needed to understand how viewers interpret the difference between the random and the actual attributions.