Goto

Collaborating Authors

 Deep Learning


What's the deal with deep learning?

#artificialintelligence

Facial recognition: controversial as it stands right now, facial recognition is still being introduced in a multitude of services and applications. Probably the most renowned one is Facebook's use of deep-learning-based recognition to tag (News - Alert) images uploaded to the platform. Some of the tests that are being conducted right now have facial recognition applications as part of security systems in public spaces. The purpose is to use neural networks to aid in the search of missing persons, while also quickly identify criminals and terrorists. In this case, deep learning is more than just the comparison of a face against a database, because the algorithms are capable of factoring in changes in hairstyles, minor surgeries, and even modifications due to the conditions of the place where the image is taken.


Nvidia creates the world's first video game demo using AI generated graphics โ€“ Fanatical Futurist by International Keynote Speaker Matthew Griffin

#artificialintelligence

Connect, download a free E-Book, watch a keynote, or browse my blog. The recent boom in Artificial Intelligence (AI) has led to the emergence of an entirely new field of research dedicated to using AI's, or Creative Machines as they're also known, to create synthetic content โ€“ or in laymans terms create "fake" digital content without the involvement of humans that includes everything from audio tracks, and imagery, to videos. As part of this new trend I recently reported how Promethean AI was using its AI to help people create game environments just by talking and describing what they wanted, and how Nvidia had created an AI that could take in a real video feed, from a city for example, and transform it in real time into digital content that could be used to create game environments as well as VR worlds. And at the time both of these were huge breakthroughs in the field. Now, in their latest research the same team behind the original Nvidia breakthrough have published research showing how AI generated video and visuals can be combined with a traditional video game engine "to create a hybrid graphics system" that could one day be used in video games, movies, and virtual reality.


Opinion Ma vs Musk on the supremacy of machines over mankind

#artificialintelligence

It was a rambling discussion in Shanghai last week that brought Artificial Intelligence into focus once again. In this case, it was the clearly jet-lagged Elon Musk, a technology legend and disruptor, and a very composed Jack Ma, the largest purveyor of goods over the internet. While the discussion rambled all over, from Pong to The Matrix to Mars, it was Musk shooting down Ma's optimism by saying that these could be "famous last words", that caught my attention. Musk has been consistently pessimistic about AI. In another famous joust, this time with Facebook's Mark Zuckerberg, he tweeted that "his (Zuckerberg's) understanding of the subject is pretty limited".


Independent Subspace Analysis for Unsupervised Learning of Disentangled Representations

arXiv.org Machine Learning

Recently there has been an increased interest in unsupervised learning of disentangled representations using the Variational Autoencoder (VAE) framework. Most of the existing work has focused largely on modifying the variational cost function to achieve this goal. We first show that these modifications, e.g. beta-VAE, simplify the tendency of variational inference to underfit causing pathological over-pruning and over-orthogonalization of learned components. Second we propose a complementary approach: to modify the probabilistic model with a structured latent prior. This prior allows to discover latent variable representations that are structured into a hierarchy of independent vector spaces. The proposed prior has three major advantages: First, in contrast to the standard VAE normal prior the proposed prior is not rotationally invariant. This resolves the problem of unidentifiability of the standard VAE normal prior. Second, we demonstrate that the proposed prior encourages a disentangled latent representation which facilitates learning of disentangled representations. Third, extensive quantitative experiments demonstrate that the prior significantly mitigates the trade-off between reconstruction loss and disentanglement over the state of the art.


Intensity augmentation for domain transfer of whole breast segmentation in MRI

arXiv.org Machine Learning

The segmentation of the breast from the chest wall is an important first step in the analysis of breast magnetic resonance images. 3D U-nets have been shown to obtain high segmentation accuracy and appear to generalize well when trained on one scanner type and tested on another scanner, provided that a very similar T1-weighted MR protocol is used. There has, however, been little work addressing the problem of domain adaptation when image intensities or patient orientation differ markedly between the training set and an unseen test set. To overcome the domain shift we propose to apply extensive intensity augmentation in addition to geometric augmentation during training. We explored both style transfer and a novel intensity remapping approach as intensity augmentation strategies. For our experiments, we trained a 3D U-net on T1-weighted scans and tested on T2-weighted scans. By applying intensity augmentation we increased segmentation performance from a DSC of 0.71 to 0.90. This performance is very close to the baseline performance of training and testing on T2-weighted scans (0.92). Furthermore, we applied our network to an independent test set made up of publicly available scans acquired using a T1-weighted TWIST sequence and a different coil configuration. On this dataset we obtained a performance of 0.89, close to the inter-observer variability of the ground truth segmentations (0.92). Our results show that using intensity augmentation in addition to geometric augmentation is a suitable method to overcome the intensity domain shift and we expect it to be useful for a wide range of segmentation tasks.


Hyper-Graph-Network Decoders for Block Codes

arXiv.org Machine Learning

Neural decoders were shown to outperform classical message passing techniques for short BCH codes. In this work, we extend these results to much larger families of algebraic block codes, by performing message passing with graph neural networks. The parameters of the sub-network at each variable-node in the Tanner graph are obtained from a hypernetwork that receives the absolute values of the current message as input. To add stability, we employ a simplified version of the arctanh activation that is based on a high order Taylor approximation of this activation function. Our results show that for a large number of algebraic block codes, from diverse families of codes (BCH, LDPC, Polar), the decoding obtained with our method outperforms the vanilla belief propagation method as well as other learning techniques from the literature.


DeepIST: Deep Image-based Spatio-Temporal Network for Travel Time Estimation

arXiv.org Machine Learning

Estimating the travel time for a given path is a fundamental problem in many urban transportation systems. However, prior works fail to well capture moving behaviors embedded in paths and thus do not estimate the travel time accurately. To fill in this gap, in this work, we propose a novel neural network framework, namely {\em Deep Image-based Spatio-Temporal network (DeepIST)}, for travel time estimation of a given path. The novelty of DeepIST lies in the following aspects: 1) we propose to plot a path as a sequence of "generalized images" which include sub-paths along with additional information, such as traffic conditions, road network and traffic signals, in order to harness the power of convolutional neural network model (CNN) on image processing; 2) we design a novel two-dimensional CNN, namely {\em PathCNN}, to extract spatial patterns for lines in images by regularization and adopting multiple pooling methods; and 3) we apply a one-dimensional CNN to capture temporal patterns among the spatial patterns along the paths for the estimation. Empirical results show that DeepIST soundly outperforms the state-of-the-art travel time estimation models by 24.37\% to 25.64\% of mean absolute error (MAE) in multiple large-scale real-world datasets.


Doppler Invariant Demodulation for Shallow Water Acoustic Communications Using Deep Belief Networks

arXiv.org Machine Learning

--Shallow water environments create a challenging channel for communications. In this paper, we focus on the challenges posed by the frequency-selective signal distortion called the Doppler effect. We explore the design and performance of machine learning (ML) based demodulation methods -- (1) Deep Belief Network-feed forward Neural Network (DBN-NN) and (2) Deep Belief Network-Convolutional Neural Network (DBN-CNN) in the physical layer of Shallow Water Acoustic Communication (SW AC). The proposed method comprises of a ML based feature extraction method and classification technique. First, the feature extraction converts the received signals to feature images. An analysis of the ML based proposed demodulation shows that despite the presence of instantaneous frequencies, the performance of the algorithm shows an invariance with a small 2dB error margin in terms of bit error rate (BER).


Understanding ML driven HPC: Applications and Infrastructure

arXiv.org Machine Learning

We recently outlined the vision of "Learning Everywhere" which captures the possibility and impact of how learning methods and traditional HPC methods can be coupled together. A primary driver of such coupling is the promise that Machine Learning (ML) will give major performance improvements for traditional HPC simulations. Motivated by this potential, the ML around HPC class of integration is of particular significance. In a related follow-up paper, we provided an initial taxonomy for integrating learning around HPC methods. In this paper, which is part of the Learning Everywhere series, we discuss "how" learning methods and HPC simulations are being integrated to enhance effective performance of computations. This paper identifies several modes --- substitution, assimilation, and control, in which learning methods integrate with HPC simulations and provide representative applications in each mode. This paper discusses some open research questions and we hope will motivate and clear the ground for MLaroundHPC benchmarks.


Receptive-field-regularized CNN variants for acoustic scene classification

arXiv.org Machine Learning

Acoustic scene classification and related tasks have been dominated by Convolutional Neural Networks (CNNs). Top-performing CNNs use mainly audio spectograms as input and borrow their architectural design primarily from computer vision. A recent study has shown that restricting the receptive field (RF) of CNNs in appropriate ways is crucial for their performance, robustness and generalization in audio tasks. One side effect of restricting the RF of CNNs is that more frequency information is lost. In this paper, we perform a systematic investigation of different RF configuration for various CNN architectures on the DCASE 2019 Task 1.A dataset. Second, we introduce Frequency Aware CNNs to compensate for the lack of frequency information caused by the restricted RF, and experimentally determine if and in what RF ranges they yield additional improvement. The result of these investigations are several well-performing submissions to different tasks in the DCASE 2019 Challenge.