Deep Learning
fastMRI: An Open Dataset and Benchmarks for Accelerated MRI
Zbontar, Jure, Knoll, Florian, Sriram, Anuroop, Muckley, Matthew J., Bruno, Mary, Defazio, Aaron, Parente, Marc, Geras, Krzysztof J., Katsnelson, Joe, Chandarana, Hersh, Zhang, Zizhao, Drozdzal, Michal, Romero, Adriana, Rabbat, Michael, Vincent, Pascal, Pinkerton, James, Wang, Duo, Yakubova, Nafissa, Owens, Erich, Zitnick, C. Lawrence, Recht, Michael P., Sodickson, Daniel K., Lui, Yvonne W.
Accelerating Magnetic Resonance Imaging (MRI) by taking fewer measurements has the potential to reduce medical costs, minimize stress to patients and make MRI possible in applications where it is currently prohibitively slow or expensive. We introduce the fastMRI dataset, a large-scale collection of both raw MR measurements and clinical MR images, that can be used for training and evaluation of machine-learning approaches to MR image reconstruction. By introducing standardized evaluation criteria and a freely-accessible dataset, our goal is to help the community make rapid advances in the state of the art for MR image reconstruction. We also provide a self-contained introduction to MRI for machine learning researchers with no medical imaging background.
Regularizing by the Variance of the Activations' Sample-Variances
Normalization techniques play an important role in supporting efficient and often more effective training of deep neural networks. While conventional methods explicitly normalize the activations, we suggest to add a loss term instead. This new loss term encourages the variance of the activations to be stable and not vary from one random mini-batch to the next. As we prove, this encourages the activations to be distributed around a few distinct modes. We also show that if the inputs are from a mixture of two Gaussians, the new loss would either join the two together, or separate between them optimally in the LDA sense, depending on the prior probabilities. Finally, we are able to link the new regularization term to the batchnorm method, which provides it with a regularization perspective. Our experiments demonstrate an improvement in accuracy over the batchnorm technique for both CNNs and fully connected networks.
Graph Refinement based Tree Extraction using Mean-Field Networks and Graph Neural Networks
Selvan, Raghavendra, Kipf, Thomas, Welling, Max, Pedersen, Jesper H, Petersen, Jens, de Bruijne, Marleen
Graph refinement, or the task of obtaining subgraphs of interest from over-complete graphs, can have many varied applications. In this work, we extract tree structures from image data by, first deriving a graph-based representation of the volumetric data and then, posing tree extraction as a graph refinement task. We present two methods to perform graph refinement. First, we use mean-field approximation (MFA) to approximate the posterior density over the subgraphs from which the optimal subgraph of interest can be estimated. Mean field networks (MFNs) are used for inference based on the interpretation that iterations of MFA can be seen as feed-forward operations in a neural network. This allows us to learn the model parameters using gradient descent. Second, we present a supervised learning approach using graph neural networks (GNNs) which can be seen as generalisations of MFNs. Subgraphs are obtained by jointly training a GNN based encoder-decoder pair, wherein the encoder learns useful edge embeddings from which the edge probabilities are predicted using a simple decoder. We discuss connections between the two classes of methods and compare them for the task of extracting airways from 3D, low-dose, chest CT data. We show that both the MFN and GNN models show significant improvement when compared to a baseline method, that is similar to a top performing method in the EXACT'09 Challenge, in detecting more branches.
Learning from Multiview Correlations in Open-Domain Videos
Holzenberger, Nils, Palaskar, Shruti, Madhyastha, Pranava, Metze, Florian, Arora, Raman
An increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views to learn better representations. This is further complicated by the existence of a latent alignment between views, such as between speech and its transcription, and by the multitude of choices for the learning objective. We explore an advanced, correlation-based representation learning method on a 4-way parallel, multimodal dataset, and assess the quality of the learned representations on retrieval-based tasks. We show that the proposed approach produces rich representations that capture most of the information shared across views. Our best models for speech and textual modalities achieve retrieval rates from 70.7% to 96.9% on open-domain, user-generated instructional videos. This shows it is possible to learn reliable representations across disparate, unaligned and noisy modalities, and encourages using the proposed approach on larger datasets.
Structure-Based Networks for Drug Validation
Cangea, Cฤtฤlina, Grauslys, Arturas, Liรฒ, Pietro, Falciani, Francesco
Classifying chemicals according to putative modes of action (MOAs) is of paramount importance in the context of risk assessment. However, current methods are only able to handle a very small proportion of the existing chemicals. We address this issue by proposing an integrative deep learning architecture that learns a joint representation from molecular structures of drugs and their effects on human cells. Our choice of architecture is motivated by the significant influence of a drug's chemical structure on its MOA. We improve on the strong ability of a unimodal architecture (F1 score of 0.803) to classify drugs by their toxic MOAs (Verhaar scheme) through adding another learning stream that processes transcriptional responses of human cells affected by drugs. Our integrative model achieves an even higher classification performance on the LINCS L1000 dataset - the error is reduced by 4.6%. We believe that our method can be used to extend the current Verhaar scheme and constitute a basis for fast drug validation and risk assessment.
Resource Mention Extraction for MOOC Discussion Forums
An, Ya-Hui, Pan, Liangming, Kan, Min-Yen, Dong, Qiang, Fu, Yan
In discussions hosted on discussion forums for Massive Online Open Courses (MOOCs), references to online learning resources are often of central importance. However they are usually mentioned in free text, without appropriate hyperlinking to their associated resource. Automated learning resource mention hyperlinking and categorization will facilitate discussion and searching within MOOC forums, and also benefit the contextualization of such resources across disparate views. We propose the novel problem of learning resource mention identification inMOOC forums; i.e., to identify resource mentions in discussions, and classify them into predefined resource types. As this is a novel task with no publicly available data, we first contribute a large-scale labeled dataset - dubbed the Forum Resource Mention (FoRM) dataset - to facilitate our current research and future research on this task. FoRM contains over 10, 000 real-world forum threads in collaboration with Coursera, with more than 23, 000 manually labeled resource mentions. We then formulate this task as a sequence tagging problem and investigate solutionarchitectures to address the problem. Corresponding author Email address: peterpan10211020@gmail.com (Liangming Pan) Preprint submitted to Elsevier November 22, 2018 two major challenges that hinder the application of sequence tagging models tothe task: (1) the diversity of resource mention expression, and (2) long-range contextual dependencies. We address these challenges by incorporating character-leveland thread context information into a LSTM-CRF model. First, we incorporate a character encoder to address the out-ofvocabulary problemcaused by the diversity of mention expressions. Second, to address the context dependency challenge, we encode thread contexts using anRNN-based context encoder, and apply the attention mechanism to selectively leverage useful context information during sequence tagging. Experiments onFoRM show that the proposed method improves the baseline deep sequence tagging models notably, significantly bettering performance on instances that exemplify the two challenges.
Artificial Intelligence-Defined 5G Radio Access Networks
Yao, Miao, Sohul, Munawwar, Marojevic, Vuk, Reed, Jeffrey H.
ABSTRACT Massive multiple-input multiple-output antenna systems, millimeter wave communications, and ultra-dense networks have been widely perceived as the three key enablers that facilitate the development and deployment of 5G systems. This article discusses the intelligent agent that combines sensing, learning, and optimizing to facilitate these enablers. We present a flexible, rapidly deployable, and cross-layer artificial intelligence (AI)-based framework to enable the imminent and future demands on 5G and beyond. We present example AIenabled 5G use cases that accommodate important 5G-specific capabilities and discuss the value of AI for enabling network evolution. I. Introduction Does 5G cellular communications technology in the age of intelligence really look like the Thomas W. Lawson Schooner (the last of the large cargo sailing ships) of modern times? However, concerns are raised whether this is a revolutionary leap from today's wireless communications or a simple piling upof less innovative wireless functionalities. The International TelecommunicationUnion (ITU) classifies 5G into three categories of usage scenarios: enhanced mobile broadband (eMBB),massive machine-type communication (mMTC), and ultra-reliable and low latency communication (URLLC) to account for more diverse services and resourcehungry applications.eMBB is a service category that addresses bandwidth-hungryapplications, such as massive video streaming and virtual/augmented reality (VR/AR). URLLC is a service category that supports latency sensitive services including autonomous driving,drones and the tactile Internet.
Machine Learning & Deep Learning
Starting off, you'll learn about Artificial Intelligence and then move to machine learning and deep learning. You will further learn how machine learning is different from deep learning, the various kinds of algorithms that fall under these two domains of learning. Finally, you will be introduced to some real-life applications where machine learning and deep learning is being applied. Well as the name suggests, artificial intelligence commonly known as AI is a way of artificially making a computer intelligent. Can you think of a computer or maybe a robot which can do various tasks similar to humans?
How the Softmax Output is Misleading for Evaluating the Strength of Adversarial Examples
Ozbulak, Utku, De Neve, Wesley, Van Messem, Arnout
Even before deep learning architectures became the de facto models for complex computer vision tasks, the softmax function was, given its elegant properties, already used to analyze the predictions of feedforward neural networks. Nowadays, the output of the softmax function is also commonly used to assess the strength of adversarial examples: malicious data points designed to fail machine learning models during the testing phase. However, in this paper, we show that it is possible to generate adversarial examples that take advantage of some properties of the softmax function, leading to undesired outcomes when interpreting the strength of the adversarial examples at hand. Specifically, we argue that the output of the softmax function is a poor indicator when the strength of an adversarial example is analyzed and that this indicator can be easily tricked by already existing methods for adversarial example generation.
Seeing in the dark with recurrent convolutional neural networks
Classical convolutional neural networks (cCNNs) are very good at categorizing objects in images. But, unlike human vision which is relatively robust to noise in images, the performance of cCNNs declines quickly as image quality worsens. Here we propose to use recurrent connections within the convolutional layers to make networks robust against pixel noise such as could arise from imaging at low light levels, and thereby significantly increase their performance when tested with simulated noisy video sequences. We show that cCNNs classify images with high signal to noise ratios (SNRs) well, but are easily outperformed when tested with low SNR images (high noise levels) by convolutional neural networks that have recurrency added to convolutional layers, henceforth referred to as gruCNNs. Addition of Bayes-optimal temporal integration to allow the cCNN to integrate multiple image frames still does not match gruCNN performance. Additionally, we show that at low SNRs, the probabilities predicted by the gruCNN (after calibration) have higher confidence than those predicted by the cCNN. We propose to consider recurrent connections in the early stages of neural networks as a solution to computer vision under imperfect lighting conditions and noisy environments; challenges faced during real-time video streams of autonomous driving at night, during rain or snow, and other non-ideal situations.