Goto

Collaborating Authors

 Asia


Gradient Acceleration in Activation Functions

arXiv.org Machine Learning

Dropout has been one of standard approaches to train deep neural networks, and it is known to regularize large models to avoid overfitting. The effect of dropout has been explained by avoiding co-adaptation. In this paper, however, we propose a new explanation of why dropout works and propose a new technique to design better activation functions. First, we show that dropout is an optimization technique to push the input towards the saturation area of nonlinear activation function by accelerating gradient information flowing even in the saturation area in backpropagation. Based on this explanation, we propose a new technique for activation functions, gradient acceleration in activation function (GAAF), that accelerates gradients to flow even in the saturation area. Then, input to the activation function can climb onto the saturation area which makes the network more robust because the model converges on a flat region. Experiment results support our explanation of dropout and confirm that the proposed GAAF technique improves performances with expected properties.


On the Implicit Bias of Dropout

arXiv.org Artificial Intelligence

Algorithmic approaches endow deep learning systems with implicit bias that helps them generalize even in over-parametrized settings. In this paper, we focus on understanding such a bias induced in learning through dropout, a popular technique to avoid overfitting in deep learning. For single hidden-layer linear neural networks, we show that dropout tends to make the norm of incoming/outgoing weight vectors of all the hidden nodes equal. In addition, we provide a complete characterization of the optimization landscape induced by dropout.


Cycle Consistent Adversarial Denoising Network for Multiphase Coronary CT Angiography

arXiv.org Artificial Intelligence

Abstract--In coronary CT angiography, a series of CT images are taken at different levels of radiation dose during the examination. Although this reduces the total radiation dose, the image quality during the low-dose phases is significantly degraded. T o address this problem, here we propose a novel semi-supervised learning technique that can remove the noises of the CT images obtained in the low-dose phases by learning from the CT images in the routine dose phases. Although a supervised learning approach is not possible due to the differences in the underlying heart structure in two phases, the images in the two phases are closely related so that we propose a cycle-consistent adversarial denoising network to learn the non-degenerate mapping between the low and high dose cardiac phases. Experimental results showed that the proposed method effectively reduces the noise in the low-dose CT image while the preserving detailed texture and edge information. Moreover, thanks to the cyclic consistency and identity loss, the proposed network does not create any artificial features that are not present in the input images. Visual grading and quality evaluation also confirm that the proposed method provides significant improvement in diagnostic quality.


Identifiability of Gaussian Structural Equation Models with Dependent Errors Having Equal Variances

arXiv.org Artificial Intelligence

In this paper, we prove that some Gaussian structural equation models with dependent errors having equal variances are identifiable from their corresponding Gaussian distributions. Specifically, we prove identifiability for the Gaussian structural equation models that can be represented as Andersson-Madigan-Perlman chain graphs (Andersson et al., 2001). These chain graphs were originally developed to represent independence models. However, they are also suitable for representing causal models with additive noise (Pe\~{n}a, 2016. Our result implies then that these causal models can be identified from observational data alone. Our result generalizes the result by Peters and B\"{u}hlmann (2014), who considered independent errors having equal variances. The suitability of the equal error variances assumption should be assessed on a per domain basis.


Representation Learning on Graphs with Jumping Knowledge Networks

arXiv.org Artificial Intelligence

Recent deep learning approaches for representation learning on graphs follow a neighborhood aggregation procedure. We analyze some important properties of these models, and propose a strategy to overcome those. In particular, the range of "neighboring" nodes that a node's representation draws from strongly depends on the graph structure, analogous to the spread of a random walk. To adapt to local neighborhood properties and tasks, we explore an architecture -- jumping knowledge (JK) networks -- that flexibly leverages, for each node, different neighborhood ranges to enable better structure-aware representation. In a number of experiments on social, bioinformatics and citation networks, we demonstrate that our model achieves state-of-the-art performance. Furthermore, combining the JK framework with models like Graph Convolutional Networks, GraphSAGE and Graph Attention Networks consistently improves those models' performance.


The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems

arXiv.org Artificial Intelligence

In this paper, we present a database of emotional speech intended to be open-sourced and used for synthesis and generation purpose. It contains data for male and female actors in English and a male actor in French. The database covers 5 emotion classes so it could be suitable to build synthesis and voice transformation systems with the potential to control the emotional dimension in a continuous way. We show the data's efficiency by building a simple MLP system converting neutral to angry speech style and evaluate it via a CMOS perception test. Even though the system is a very simple one, the test show the efficiency of the data which is promising for future work.


Regularized Evolution for Image Classifier Architecture Search

arXiv.org Artificial Intelligence

The effort devoted to hand-crafting image classifiers has motivated the use of architecture search to discover them automatically. Although evolutionary algorithms have been repeatedly applied to architecture search, the architectures thus discovered have remained inferior to human-crafted ones. Here we show for the first time that artificially-evolved architectures can match or surpass human-crafted and RL-designed image classifiers. In particular, our models---named AmoebaNets---achieved a state-of-the-art accuracy of 97.87% on CIFAR-10 and top-1 accuracy of 83.1% on ImageNet. Among mobile-size models, an AmoebaNet with only 5.1M parameters also achieved a state-of-the-art top-1 accuracy of 75.1% on ImageNet. We also compared this method against strong baselines. Finally, we performed platform-aware architecture search with evolution to find a model that trains quickly on Google Cloud TPUs. This method produced an AmoebaNet that won the Stanford DAWNBench competition for lowest ImageNet training cost.


CASP Solutions for Planning in Hybrid Domains

arXiv.org Artificial Intelligence

CASP is an extension of ASP that allows for numerical constraints to be added in the rules. PDDL+ is an extension of the PDDL standard language of automated planning for modeling mixed discrete-continuous dynamics. In this paper, we present CASP solutions for dealing with PDDL+ problems, i.e., encoding from PDDL+ to CASP, and extensions to the algorithm of the EZCSP CASP solver in order to solve CASP programs arising from PDDL+ domains. An experimental analysis, performed on well-known linear and non-linear variants of PDDL+ domains, involving various configurations of the EZCSP solver, other CASP solvers, and PDDL+ planners, shows the viability of our solution.


Wake up to AI opportunity

#artificialintelligence

Niti Aayog's strategy paper on Artificial Intelligence (AI), which outlines the opportunities offered by this rising technology and the challenges posed by it, calls for special attention from governments and policymakers. The endeavour of Artificial Intelligence is the development of intelligence in machines, either by feeding into them capability for specialised tasks or, increasingly, giving them the capability, such as sensors and large amounts of data, to learn on their own, called Machine Learning. The Niti Aayog paper presents the need to create a favourable AI ecosystem in India and identifies the sectors that need to be focused on, like agriculture, education, healthcare, infrastructure and transport. The economic rewards and social benefits will be great, as will be the disruptions caused by AI. The paper estimates that AI might account for about $1 trillion, or about 15% of GDP, by 2035.


This Week's Top Stocks FB, DDD, AMZN, & TWTR Stock Forecasts Quantifying Uncertainty and Bayesian Inference

#artificialintelligence

The U.S. cotton market has remained stable since its spike in 2011, when China executed its cotton reserving and fiber hoarding plan. It is believed that U.S. cotton demand and price were artificially kept low because there are always worries that China would unexpectedly unleash its cotton stockpile, about half of the global storage. However, U.S. cotton price finally showed a revival in recent days. The ICE July cotton futures closed at 95.21 cents a pound on Tuesday, June 12, the highest level for a front-month future contract in the last 6 years. The revival could be attributed to multiple factors, with an emphasis on the worries about insufficient rain in the cotton-growing areas and the newly issued import quotas from China.