Deep Learning
Towards Sparse Hierarchical Graph Classifiers
Cangea, Cătălina, Veličković, Petar, Jovanović, Nikola, Kipf, Thomas, Liò, Pietro
Recent advances in representation learning on graphs, mainly leveraging graph convolutional networks, have brought a substantial improvement on many graph-based benchmark tasks. While novel approaches to learning node embeddings are highly suitable for node classification and link prediction, their application to graph classification (predicting a single label for the entire graph) remains mostly rudimentary, typically using a single global pooling step to aggregate node features or a hand-designed, fixed heuristic for hierarchical coarsening of the graph structure. An important step towards ameliorating this is differentiable graph coarsening---the ability to reduce the size of the graph in an adaptive, data-dependent manner within a graph neural network pipeline, analogous to image downsampling within CNNs. However, the previous prominent approach to pooling has quadratic memory requirements during training and is therefore not scalable to large graphs. Here we combine several recent advances in graph neural network design to demonstrate that competitive hierarchical graph classification results are possible without sacrificing sparsity. Our results are verified on several established graph classification benchmarks, and highlight an important direction for future research in graph-based neural networks.
Relation Mention Extraction from Noisy Data with Hierarchical Reinforcement Learning
Feng, Jun, Huang, Minlie, Zhang, Yijie, Yang, Yang, Zhu, Xiaoyan
In this paper we address a task of relation mention extraction from noisy data: extracting representative phrases for a particular relation from noisy sentences that are collected via distant supervision. Despite its significance and value in many downstream applications, this task is less studied on noisy data. The major challenges exists in 1) the lack of annotation on mention phrases, and more severely, 2) handling noisy sentences which do not express a relation at all. To address the two challenges, we formulate the task as a semi-Markov decision process and propose a novel hierarchical reinforcement learning model. Our model consists of a top-level sentence selector to remove noisy sentences, a low-level mention extractor to extract relation mentions, and a reward estimator to provide signals to guide data denoising and mention extraction without explicit annotations. Experimental results show that our model is effective to extract relation mentions from noisy data.
Training neural audio classifiers with few data
Pons, Jordi, Serrà, Joan, Serra, Xavier
These studies are mostly based on publiclyavailable datasets, where each class typically contains more than 100 audio examples [5, 6, 7, 8, 9]. Contrastingly, only few works study the problem of training neural audio classifiers with few audio examples (for instance, less than 10 per class) [10, 11, 12, 13]. In this work, we study how a number of neural network architectures perform in such situation. Two primary reasons motivate our work: (i) given that humans are able to learn novel concepts from few examples, we aim to quantify up to what extent such behavior is possible in current neural machine listening systems; and (ii) provided that data curation processes are tedious and expensive, it is unreasonable to assume that sizable amounts of annotated audio are always available for training neural network classifiers. The challenge of training neural networks with few audio data has been previously addressed. For example, Morfi and Stowell [12] approached the problem via factorising an audio transcription task into two intermediate sub-tasks: event and tag detection.
How teaching AI to be curious helps machines learn for themselves
When playing a video game, what motivates you to carry on? This question is perhaps too broad to yield a single answer, but if you had to sum up why you accept that next quest, jump into a new level, or cave and play just one more turn, the simplest explanation might be "curiosity" -- just to see what happens next. And as it turns out, curiosity is a very effective motivator when teaching AI to play video games, too. Research published this week by artificial intelligence lab OpenAI explains how an AI agent with a sense of curiosity outperformed its predecessors playing the classic 1984 Atari game Montezuma's Revenge. Becoming skilled at Montezuma's Revenge is not a milestone equivalent to beating Go or Dota 2, but it's still a notable advance.
A Deep Learning Machine On Azure From The App Marketplace
I've run a lot of machine learning/A.I. projects as toys, and even a few not very complex ones in production. Normally they run on the CPU, and in only one instance did I use a GPU ... the projects simply didn't require it. Sometimes however, you come across something you need to try out, and it needs a STONKIN BIG MOTHA of a machine to really get its teeth stuck in. I had to do that recently and found the quickest way to get started was to spin up what I needed using the pre-configured Azure Deep Learning environment, then drop it when I was finished. This article walks through the process that is actually rather pleasantly simple.
IBM, Harvard develop tool to tackle black box problem in AI translation
In recent years, machine translation has improved immensely thanks to advances in deep learning and neural networks. However, the advantages of neural networks come at the cost of not knowing for sure what goes on inside them, which means it's hard to troubleshoot their mistakes, such as when they translate "good morning" in Arabic to "attack them" in Hebrew. Researchers at IBM and Harvard University have developed a new debugging tool to address this issue. Presented at the IEEE Conference on Visual Analytics Science and Technology in Berlin last week, the tool lets creators of deep learning applications visualize the decision-making an AI makes when translating a sequence of words from one language to another. Called Seq2Seq-Vis, the tool is one of the several efforts that aim to interpret decisions made by deep neural networks.
Deep Learning for Breast Cancer Identification from Histopathological Images
Breast cancer is one of the leading causes of death by cancer for women. Early detection can give patients more treatment options. In order to detect signs of cancer, breast tissue from biopsies is stained to enhance the nuclei and cytoplasm for microscopic examination. Then, pathologists evaluate the extent of any abnormal structural variation to determine whether there are tumors. Since the majority of biopsies find normal and benign results, most of the manual labelling of these microscopic images is redundant.
Deep learning is not a replacement for human creativity, period
This article is part of Demystifying AI, a series of posts that (try to) disambiguate the jargon and myths surrounding AI. An AI-made portrait sold for $432,500 at a famous auction last week. This was a story that was widely discussed in tech media in the past week, with some suggesting the development marked a threat for human artists. This is just one of the many stories of progress in deep learning that triggers sensational headlines about AI manifesting artistic creativity that is on par with humans. "AI songwriting has arrived" and "AI will soon write better novels than humans" are just some of the stories that have surfaced on mainstream media in the past few months.
A Deep Learning Framework for Single-Sided Sound Speed Inversion in Medical Ultrasound
Feigin, Micha, Freedman, Daniel, Anthony, Brian W.
Ultrasound elastography is gaining traction as an accessible and useful diagnostic tool for such things as cancer detection and differentiation as well as liver and thyroid disease diagnostics. Unfortunately, state of the art acoustic radiation force techniques, essential to promote this goal, are limited to high end ultrasound hardware due to high power requirements; are extremely sensitive to patient and sonographer motion; and generally suffer from low frame rates. Researchers have shown that pressure wave velocity possesses similar diagnostic abilities to shear wave velocity. Using pressure waves removes the need for generating shear waves, which in turn enables elasticity based diagnostic techniques on portable and low cost devices. However, current travel time tomography and full waveform inversion techniques for recovering pressure wave velocities require a full circumferential field of view. Focus based techniques, on the other hand, provide only localized measurements, are sensitive to the intermediate medium and require capturing multiple frames. In this paper, we present a single sided sound speed inversion solution using a fully convolutional deep neural network. We show that it is possible to invert for longitudinal sound speed in soft tissue at real time frame rates. For the computation, analysis is performed on channel data information from three diagonal plane waves. This is the first step towards a full waveform solver using a Deep Learning framework for the elastic and viscoelastic inverse problem.
Frequentist uncertainty estimates for deep learning
Tagasovska, Natasa, Lopez-Paz, David
We provide frequentist estimates of aleatoric and epistemic uncertainty for deep neural networks. To estimate aleatoric uncertainty we propose simultaneous quantile regression, a loss function to learn all the conditional quantiles of a given target variable. These quantiles lead to well-calibrated prediction intervals. To estimate epistemic uncertainty we propose training certificates, a collection of diverse non-trivial functions that map all training samples to zero. These certificates map out-of-distribution examples to non-zero values, signaling high epistemic uncertainty. We compare our proposals to prior art in various experiments.