Goto

Collaborating Authors

 Asia


A Deep Neural Network Surrogate for High-Dimensional Random Partial Differential Equations

arXiv.org Machine Learning

Developing efficient numerical algorithms for high dimensional random Partial Differential Equations (PDEs) has been a challenging task due to the well-known curse of dimensionality problem. We present a new framework for solving high-dimensional PDEs characterized by random parameters based on a deep learning approach. The random PDE is approximated by a feed-forward fully-connected deep neural network, with either strong or weak enforcement of initial and boundary constraints. The framework is mesh-free, and can handle irregular computational domains. Parameters of the approximating deep neural network are determined iteratively using variants of the Stochastic Gradient Descent (SGD) algorithm, which removes the memory issues that some existing algorithms for random PDEs are currently experiencing. The performance of the proposed framework in accurate estimation of solution to random PDEs is examined by implementing it for diffusion and heat conduction problems. Results are compared with the sampling-based finite element results, and suggest that the proposed framework achieves satisfactory accuracy and can handle high-dimensional random PDEs. A discussion on the advantages of the proposed method is also provided. Introduction Partial differential equations (PDEs) are used to describe a variety of physical phenomena such as fluid dynamics, quantum mechanics, and elasticity.


Using Social Network Information in Bayesian Truth Discovery

arXiv.org Machine Learning

We investigate the problem of truth discovery based on opinions from multiple agents who may be unreliable or biased. We consider the case where agents' reliabilities or biases are correlated if they belong to the same community, which defines a group of agents with similar opinions regarding a particular event. An agent can belong to different communities for different events, and these communities are unknown \emph{a priori}. We incorporate knowledge of the agents' social network in our truth discovery framework and develop Laplace variational inference methods to estimate agents' reliabilities, communities, and the event states. We also develop a stochastic variational inference method to scale our model to large social networks. Simulations and experiments on real data suggest that when observations are sparse, our proposed methods perform better than several other inference methods, including majority voting, the popular Bayesian Classifier Combination (BCC) method, and the Community BCC method.


Probabilistic FastText for Multi-Sense Word Embeddings

arXiv.org Machine Learning

We introduce Probabilistic FastText, a new model for word embeddings that can capture multiple word senses, sub-word structure, and uncertainty information. In particular, we represent each word with a Gaussian mixture density, where the mean of a mixture component is given by the sum of n-grams. This representation allows the model to share statistical strength across sub-word structures (e.g. Latin roots), producing accurate representations of rare, misspelt, or even unseen words. Moreover, each component of the mixture can capture a different word sense. Probabilistic FastText outperforms both FastText, which has no probabilistic model, and dictionary-level probabilistic embeddings, which do not incorporate subword structures, on several word-similarity benchmarks, including English RareWord and foreign language datasets. We also achieve state-of-art performance on benchmarks that measure ability to discern different meanings. Thus, the proposed model is the first to achieve multi-sense representations while having enriched semantics on rare words.


Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization

arXiv.org Machine Learning

Recent studies have illustrated that stochastic gradient Markov Chain Monte Carlo techniques have a strong potential in non-convex optimization, where local and global convergence guarantees can be shown under certain conditions. By building up on this recent theory, in this study, we develop an asynchronous-parallel stochastic L-BFGS algorithm for non-convex optimization. The proposed algorithm is suitable for both distributed and shared-memory settings. We provide formal theoretical analysis and show that the proposed method achieves an ergodic convergence rate of ${\cal O}(1/\sqrt{N})$ ($N$ being the total number of iterations) and it can achieve a linear speedup under certain conditions. We perform several experiments on both synthetic and real datasets. The results support our theory and show that the proposed algorithm provides a significant speedup over the recently proposed synchronous distributed L-BFGS algorithm.


Dimensionality-Driven Learning with Noisy Labels

arXiv.org Machine Learning

Datasets with significant proportions of noisy (incorrect) class labels present challenges for training accurate Deep Neural Networks (DNNs). We propose a new perspective for understanding DNN generalization for such datasets, by investigating the dimensionality of the deep representation subspace of training samples. We show that from a dimensionality perspective, DNNs exhibit quite distinctive learning styles when trained with clean labels versus when trained with a proportion of noisy labels. Based on this finding, we develop a new dimensionality-driven learning strategy, which monitors the dimensionality of subspaces during training and adapts the loss function accordingly. We empirically demonstrate that our approach is highly tolerant to significant proportions of noisy labels, and can effectively learn low-dimensional local subspaces that capture the data distribution.


Exact Low Tubal Rank Tensor Recovery from Gaussian Measurements

arXiv.org Machine Learning

The recent proposed Tensor Nuclear Norm (TNN) [Lu et al., 2016; 2018a] is an interesting convex penalty induced by the tensor SVD [Kilmer and Martin, 2011]. It plays a similar role as the matrix nuclear norm which is the convex surrogate of the matrix rank. Considering that the TNN based Tensor Robust PCA [Lu et al., 2018a] is an elegant extension of Robust PCA with a similar tight recovery bound, it is natural to solve other low rank tensor recovery problems extended from the matrix cases. However, the extensions and proofs are generally tedious. The general atomic norm provides a unified view of low-complexity structures induced norms, e.g., the $\ell_1$-norm and nuclear norm. The sharp estimates of the required number of generic measurements for exact recovery based on the atomic norm are known in the literature. In this work, with a careful choice of the atomic set, we prove that TNN is a special atomic norm. Then by computing the Gaussian width of certain cone which is necessary for the sharp estimate, we achieve a simple bound for guaranteed low tubal rank tensor recovery from Gaussian measurements. Specifically, we show that by solving a TNN minimization problem, the underlying tensor of size $n_1\times n_2\times n_3$ with tubal rank $r$ can be exactly recovered when the given number of Gaussian measurements is $O(r(n_1+n_2-r)n_3)$. It is order optimal when comparing with the degrees of freedom $r(n_1+n_2-r)n_3$. Beyond the Gaussian mapping, we also give the recovery guarantee of tensor completion based on the uniform random mapping by TNN minimization. Numerical experiments verify our theoretical results.


Large scale classification in deep neural network with Label Mapping

arXiv.org Machine Learning

In recent years, deep neural network is widely used in machine learning. The multi-class classification problem is a class of important problem in machine learning. However, in order to solve those types of multi-class classification problems effectively, the required network size should have hyper-linear growth with respect to the number of classes. Therefore, it is infeasible to solve the multi-class classification problem using deep neural network when the number of classes are huge. This paper presents a method, so called Label Mapping (LM), to solve this problem by decomposing the original classification problem to several smaller sub-problems which are solvable theoretically. Our method is an ensemble method like error-correcting output codes (ECOC), but it allows base learners to be multi-class classifiers with different number of class labels. We propose two design principles for LM, one is to maximize the number of base classifier which can separate two different classes, and the other is to keep all base learners to be independent as possible in order to reduce the redundant information. Based on these principles, two different LM algorithms are derived using number theory and information theory. Since each base learner can be trained independently, it is easy to scale our method into a large scale training system. Experiments show that our proposed method outperforms the standard one-hot encoding and ECOC significantly in terms of accuracy and model complexity.


India's Largest Companies: IT Outsourcers Lead The Way

Forbes - Tech

Infosys and its local rivals in India are some the largest IT outsourcers on the planet, and definitely the biggest in any emerging market. If you include New Jersey-based Cognizant in the mix, cofounded by Lakshmi Narayanan from Tata Consultancy, then six of the biggest IT services companies in the world are run by Indians. As it stands, five of the largest IT consultancies call India home. That's more than China, more than Ireland and France, and only two less than the United States if you consider the companies whose outsourcing practice is at the core of their business, like IBM for instance. India's biggest IT outsourcer of course is Tata Consultancy Services, ranked #404 on the Forbes Global 2000 list.


The 10 Biggest Artificial Intelligence Startups in The World - Nanalyze

#artificialintelligence

The world's most powerful person used to be Vladimir Putin. This year he was defeated by the Chinese president, Xi Jinping, according to the Forbes ranking of powerful people who make you question what you've been doing with your life so far. It's safe to say that Mr. Putin knows plenty about power, and he believes that advances in artificial intelligence (AI) will not only change the world as we know it, but the global balance of power as well. It looks like other world leaders agree, with China and the US fighting for AI supremacy and the EU scrambling to catch up. Real growth is fueled by cold hard cash, so we've put together a list of the 10 biggest artificial intelligence startups in the world by funding.


Fire-fighting 'dragon' robot with the body of a hose can wiggle into windows to put out a blaze

Daily Mail - Science & tech

Japanese researchers have developed an astonishing robot with a snake-like body that is capable of fighting fires. The'dragon robot' is capable of wriggling into hard-to-reach gaps between structures and windows several floors up. It therefore can extinguish blazes traditional firefighters might not be able to reach. Researchers from Tohoku University and National Institute of Technology, Hachinohe College presented the robot at the International Conference on Robotics and Automation last month in Brisbane, Australia. The machine, called the DragonFireFighter, has the ability to lift itself off the ground and fly using high pressure jets of water.