Goto

Collaborating Authors

 Deep Learning


Sparsely Grouped Input Variables for Neural Networks

arXiv.org Machine Learning

In genomic analysis, biomarker discovery, image recognition, and other systems involving machine learning, input variables can often be organized into different groups by their source or semantic category. Eliminating some groups of variables can expedite the process of data acquisition and avoid over-fitting. Researchers have used the group lasso to ensure group sparsity in linear models and have extended it to create compact neural networks in meta-learning. Different from previous studies, we use multi-layer non-linear neural networks to find sparse groups for input variables. We propose a new loss function to regularize parameters for grouped input variables, design a new optimization algorithm for this loss function, and test these methods in three real-world settings. We achieve group sparsity for three datasets, maintaining satisfying results while excluding one nucleotide position from an RNA splicing experiment, excluding 89.9% of stimuli from an eye-tracking experiment, and excluding 60% of image rows from an experiment on the MNIST dataset.


Detecting anthropogenic cloud perturbations with deep learning

arXiv.org Machine Learning

One of the most pressing questions in climate science is that of the effect of anthropogenic aerosol on the Earth's energy balance. Aerosols provide the `seeds' on which cloud droplets form, and changes in the amount of aerosol available to a cloud can change its brightness and other physical properties such as optical thickness and spatial extent. Clouds play a critical role in moderating global temperatures and small perturbations can lead to significant amounts of cooling or warming. Uncertainty in this effect is so large it is not currently known if it is negligible, or provides a large enough cooling to largely negate present-day warming by CO2. This work uses deep convolutional neural networks to look for two particular perturbations in clouds due to anthropogenic aerosol and assess their properties and prevalence, providing valuable insights into their climatic effects.


Orthogonal Wasserstein GANs

arXiv.org Machine Learning

Wasserstein-GANs have been introduced to address the deficiencies of generative adversarial networks (GANs) regarding the problems of vanishing gradients and mode collapse during the training, leading to improved convergence behaviour and improved image quality. However, Wasserstein-GANs require the discriminator to be Lipschitz continuous. In current state-of-the-art Wasserstein-GANs this constraint is enforced via gradient norm regularization. In this paper, we demonstrate that this regularization does not encourage a broad distribution of spectral-values in the discriminator weights, hence resulting in less fidelity in the learned distribution. We therefore investigate the possibility of substituting this Lipschitz constraint with an orthogonality constraint on the weight matrices. We compare three different weight orthogonalization techniques with regards to their convergence properties, their ability to ensure the Lipschitz condition and the achieved quality of the learned distribution. In addition, we provide a comparison to Wasserstein-GANs trained with current state-of-the-art methods, where we demonstrate the potential of solely using orthogonality-based regularization. In this context, we propose an improved training procedure for Wasserstein-GANs which utilizes orthogonalization to further increase its generalization capability. Finally, we provide a novel metric to evaluate the generalization capabilities of the discriminators of different Wasserstein-GANs.


Deep Learning to Scale up Time Series Traffic Prediction

arXiv.org Machine Learning

--The transport literature is dense regarding short-term traffic predictions, up to the scale of 1 hour, yet less dense for long-term traffic predictions. The transport literature is also sparse when it comes to city-scale traffic predictions, mainly because of low data availability. The main question we try to answer in this work is to which extent the approaches used for short-term prediction at a link level can be scaled up for long-term prediction at a city scale. We investigate a city-scale traffic dataset with 14 weeks of speed observations collected every 15 minutes over 1098 segments in the hypercenter of Los Angeles, California. We look at a variety of machine learning and deep learning predictors for link-based predictions, and investigate ways to make such predictors scale up for larger areas, with brute force, clustering, and model design approaches. In particular we propose a novel deep learning spatiotemporal predictor inspired from recent works on recommender systems. We discuss the potential of including spatiotemporal features into the predictors, and conclude that modelling such features can be helpful for long-term predictions, while simpler predictors achieve very satisfactory performance for link-based and short-term forecasting. The tradeoff is discussed not only in terms of prediction accuracy vs prediction horizon but also in terms of training time and model sizing. Traffic prediction in urban transport networks is a central task for the real-time operation of transportation systems, such as route planning, route guidance, on-demand mobility services Simonetto et al. (2019). In principle this task can be achieved with the help of an increasing large volume of observed traffic data that can be made available through, e.g., on-road sensors, GPS data, cameras, social media Zhu et al. (2019). In reality, the access to such data is limited as big traffic data sets are generally owned by specific companies and deemed as proprietary information and a valuable source of business.


Deep Networks with Adaptive Nystr\"om Approximation

arXiv.org Machine Learning

Recent work has focused on combining kernel methods and deep learning to exploit the best of the two approaches. Here, we introduce a new architecture of neural networks in which we replace the top dense layers of standard convolutional architectures with an approximation of a kernel function by relying on the Nystr{\"o}m approximation. Our approach is easy and highly flexible. It is compatible with any kernel function and it allows exploiting multiple kernels. We show that our architecture has the same performance than standard architecture on datasets like SVHN and CIFAR100. One benefit of the method lies in its limited number of learnable parameters which makes it particularly suited for small training set sizes, e.g. from 5 to 20 samples per class.


Towards Oracle Knowledge Distillation with Neural Architecture Search

arXiv.org Machine Learning

We present a novel framework of knowledge distillation that is capable of learning powerful and efficient student models from ensemble teacher networks. Our approach addresses the inherent model capacity issue between teacher and student and aims to maximize benefit from teacher models during distillation by reducing their capacity gap. Specifically, we employ a neural architecture search technique to augment useful structures and operations, where the searched network is appropriate for knowledge distillation towards student models and free from sacrificing its performance by fixing the network capacity. We also introduce an oracle knowledge distillation loss to facilitate model search and distillation using an ensemble-based teacher model, where a student network is learned to imitate oracle performance of the teacher. We perform extensive experiments on the image classification datasets---CIFAR-100 and TinyImageNet---using various networks. We also show that searching for a new student model is effective in both accuracy and memory size and that the searched models often outperform their teacher models thanks to neural architecture search with oracle knowledge distillation.


Foodvisor raises $4.5 million to track what you eat using AI โ€“ TechCrunch

#artificialintelligence

French startup Foodvisor has raised a $4.5 million funding round after generating 2 million app downloads. Agrinnovation is leading the round and various business angels are also participating. I covered Foodvisor last month, so I'm not going to describe the app once again. In a few words, the startup uses deep learning to enable image recognition to detect what you're about to eat. It can detect the type of food and it also tries to estimate the weight of each item.


Open Source Projects by Google, Uber and Facebook for Data Science and AI - KDnuggets

#artificialintelligence

Open source is becoming the standard for sharing and improving technology. Some of the largest organizations in the world namely: Google, Facebook and Uber are open sourcing their own technologies that they use in their workflow to the public. This has allowed the common person to utilize technologies that are used in the biggest companies in the world. Probably the most well-known open source projects are PyTorch and Tensorflow (both coincidentally being the de-facto standard for Deep Learning).


Herring, Not Herring: Deep Learning Accelerates Detection and Classification of Underwater Species

#artificialintelligence

Canadian machine learning researchers from the University of Victoria have teamed up with government marine biologists and private remote sensing specialists to develop a system for improved detection and classification of schools of herring. The world's oceans are home to some 200,000 species of sea animals, including over 18,000 species of fish, more than 1,800 sea stars, 816 squids, 93 whales and dolphins and 8,900 clams and other bivalves, according to a 2015 report from the World Register of Marine Species. Ocean fishes come in a variety of shapes, sizes, and colors and live in many different depth and temperature environments. This diverse marine world is however under threat. A 2016 United Nations Food and Agriculture Organization's World Fisheries and Aquaculture report reveals that 89.5 percent of the world's fish stocks are either fully fished (catches are close to the maximum sustainable yield) or overfished (catches are unsustainable).


Rossum - Artificial Intelligence for data extraction from documents.

#artificialintelligence

Extract structured knowledge from an unlimited number of documents, in any format, and instantly expand your human capabilities tenfold.