Goto

Collaborating Authors

 Deep Learning


NAS-Bench-101: Towards Reproducible Neural Architecture Search

arXiv.org Machine Learning

Recent advances in neural architecture search (NAS) demand tremendous computational resources. This makes it difficult to reproduce experiments and imposes a barrier-to-entry to researchers without access to large-scale computation. We aim to ameliorate these problems by introducing NAS-Bench-101, the first public architecture dataset for NAS research. To build NAS-Bench-101, we carefully constructed a compact, yet expressive, search space, exploiting graph isomorphisms to identify 423k unique convolutional architectures. We trained and evaluated all of these architectures multiple times on CIFAR-10 and compiled the results into a large dataset. All together, NAS-Bench-101 contains the metrics of over 5 million models, the largest dataset of its kind thus far. This allows researchers to evaluate the quality of a diverse range of models in milliseconds by querying the pre-computed dataset. We demonstrate its utility by analyzing the dataset as a whole and by benchmarking a range of architecture optimization algorithms.


Short-term Road Traffic Prediction based on Deep Cluster at Large-scale Networks

arXiv.org Machine Learning

Short-term road traffic prediction (STTP) is one of the most important modules in Intelligent Transportation Systems (ITS). However, network-level STTP still remains challenging due to the difficulties both in modeling the diverse traffic patterns and tacking high-dimensional time series with low latency. Therefore, a framework combining with a deep clustering (DeepCluster) module is developed for STTP at largescale networks in this paper. The DeepCluster module is proposed to supervise the representation learning in a visualized way from the large unlabeled dataset. More specifically, to fully exploit the traffic periodicity, the raw series is first split into a number of sub-series for triplets generation. The convolutional neural networks (CNNs) with triplet loss are utilized to extract the features of shape by transferring the series into visual images. The shape-based representations are then used for road segments clustering. Thereafter, motivated by the fact that the road segments in a group have similar patterns, a model sharing strategy is further proposed to build recurrent NNs (RNNs)-based predictions through a group-based model (GM), instead of individual-based model (IM) in which one model are built for one road exclusively. Our framework can not only significantly reduce the number of models and cost, but also increase the number of training data and the diversity of samples. In the end, we evaluate the proposed framework over the network of Liuli Bridge in Beijing. Experimental results show that the DeepCluster can effectively cluster the road segments and GM can achieve comparable performance against the IM with less number of models.


The State of Sparsity in Deep Neural Networks

arXiv.org Machine Learning

We rigorously evaluate three state-of-the-art techniques for inducing sparsity in deep neural networks on two large-scale learning tasks: Transformer trained on WMT 2014 English-to-German, and ResNet-50 trained on ImageNet. Across thousands of experiments, we demonstrate that complex techniques (Molchanov et al., 2017; Louizos et al., 2017b) shown to yield high compression rates on smaller datasets perform inconsistently, and that simple magnitude pruning approaches achieve comparable or better results. Additionally, we replicate the experiments performed by (Frankle & Carbin, 2018) and (Liu et al., 2018) at scale and show that unstructured sparse architectures learned through pruning cannot be trained from scratch to the same test set performance as a model trained with joint sparsification and optimization. Together, these results highlight the need for large-scale benchmarks in the field of model compression. We open-source our code, top performing model checkpoints, and results of all hyperparameter configurations to establish rigorous baselines for future work on compression and sparsification.


Wasserstein-Wasserstein Auto-Encoders

arXiv.org Machine Learning

To address the challenges in learning deep generative models (e.g.,the blurriness of variational auto-encoder and the instability of training generative adversarial networks, we propose a novel deep generative model, named Wasserstein-Wasserstein auto-encoders (WWAE). We formulate WWAE as minimization of the penalized optimal transport between the target distribution and the generated distribution. By noticing that both the prior $P_Z$ and the aggregated posterior $Q_Z$ of the latent code Z can be well captured by Gaussians, the proposed WWAE utilizes the closed-form of the squared Wasserstein-2 distance for two Gaussians in the optimization process. As a result, WWAE does not suffer from the sampling burden and it is computationally efficient by leveraging the reparameterization trick. Numerical results evaluated on multiple benchmark datasets including MNIST, fashion- MNIST and CelebA show that WWAE learns better latent structures than VAEs and generates samples of better visual quality and higher FID scores than VAEs and GANs.


Field-aware Neural Factorization Machine for Click-Through Rate Prediction

arXiv.org Machine Learning

Recommendation systems and computing advertisements have gradually entered the field of academic research from the field of commercial applications. Click-through rate prediction is one of the core research issues because the prediction accuracy affects the user experience and the revenue of merchants and platforms. Feature engineering is very important to improve click-through rate prediction. Traditional feature engineering heavily relies on people's experience, and is difficult to construct a feature combination that can describe the complex patterns implied in the data. This paper combines traditional feature combination methods and deep neural networks to automate feature combinations to improve the accuracy of click-through rate prediction. We propose a mechannism named 'Field-aware Neural Factorization Machine' (FNFM). This model can have strong second order feature interactive learning ability like Field-aware Factorization Machine, on this basis, deep neural network is used for higher-order feature combination learning. Experiments show that the model has stronger expression ability than current deep learning feature combination models like the DeepFM, DCN and NFM.


AI researchers debate the ethics of sharing potentially harmful programs

#artificialintelligence

A recent decision by research lab OpenAI to limit the release of a new algorithm has caused controversy in the AI community. The nonprofit said it decided not to share the full version of the program, a text-generation algorithm named GPT-2, due to concerns over "malicious applications." But many AI researchers have criticized the decision, accusing the lab of exaggerating the danger posed by the work and inadvertently stoking "mass hysteria" about AI in the process. The debate has been wide-ranging and sometimes contentious. It even turned into a bit of a meme among AI researchers, who joked that they've had an amazing breakthrough in the lab, but the results were too dangerous to share at the moment.


30 Top Artificial Intelligence And Machine Learning Companies

#artificialintelligence

Artificial intelligence has become an essential part of our everyday lives. It is used in financial processes, medical examinations, logistics, publishing, and in a wide range of other fast-rising industries. According to The AI Index 2018 Annual Report by Stanford University, active AI startups in the US increased 2.1x from 2015 to 2018, while venture capital funding for US AI startups increased 4.5x from 2013 to 2017. Today, there are so many AI development companies on the market that it is becoming more and more difficult to choose the one. Based on myexperience in IT market research, I've compiled a list of best AI providers.


MultiLayerNetwork and ComputationGraph

#artificialintelligence

MultiLayerNetwork'MultiLayerNetwork' consists of a single input layer and a single output layer with a stack of layers in between them.


Recommendations for Deep Learning Neural Network Practitioners

#artificialintelligence

Deep learning neural networks are relatively straightforward to define and train given the wide adoption of open source libraries. Nevertheless, neural networks remain challenging to configure and train. In his 2012 paper titled "Practical Recommendations for Gradient-Based Training of Deep Architectures" published as a preprint and a chapter of the popular 2012 book "Neural Networks: Tricks of the Trade," Yoshua Bengio, one of the fathers of the field of deep learning, provides practical recommendations for configuring and tuning neural network models. In this post, you will step through this long and interesting paper and pick out the most relevant tips and tricks for modern deep learning practitioners. Practical Recommendations for Deep Learning Neural Network Practitioners Photo by Susanne Nilsson, some rights reserved.


Why Small Data is Essential for Advancing AI

#artificialintelligence

Everything was small data before we had big data. The scientific discoveries of the 19th and 20th centuries were all made using small data. Physicists made all calculations by hand, thus exclusively using small data. And yet, they discovered the most beautiful and most fundamental laws of nature. Moreover, they compressed them into simple rules in the form of elegant equations.