Goto

Collaborating Authors

 Africa


Nonlinear System Identification via Tensor Completion

arXiv.org Machine Learning

Function approximation from input and output data pairs constitutes a fundamental problem in supervised learning. Deep neural networks are currently the most popular method for learning to mimic the input-output relationship of a generic nonlinear system, as they have proven to be very effective in approximating complex highly nonlinear functions. In this work, we propose low-rank tensor completion as an appealing alternative for modeling and learning complex nonlinear systems. We model the interactions between the $N$ input variables and the scalar output of a system by a single N-way tensor, and setup a weighted low-rank tensor completion problem with smoothness regularization which we tackle using a block coordinate descent algorithm. We extend our method to the multi-output setting and the case of partially observed data, which cannot be readily handled by neural networks. Finally, we demonstrate the effectiveness of the approach using several regression tasks including some standard benchmarks and a challenging student grade prediction task.


Cognitive Knowledge Graph Reasoning for One-shot Relational Learning

arXiv.org Machine Learning

Inferring new facts from existing knowledge graphs (KG) with explainable reasoning processes is a significant problem and has received much attention recently. However, few studies have focused on relation types unseen in the original KG, given only one or a few instances for training. To bridge this gap, we propose CogKR for one-shot KG reasoning. The one-shot relational learning problem is tackled through two modules: the summary module summarizes the underlying relationship of the given instances, based on which the reasoning module infers the correct answers. Motivated by the dual process theory in cognitive science, in the reasoning module, a cognitive graph is built by iteratively coordinating retrieval (System 1, collecting relevant evidence intuitively) and reasoning (System 2, conducting relational reasoning over collected information). The structural information offered by the cognitive graph enables our model to aggregate pieces of evidence from multiple reasoning paths and explain the reasoning process graphically. Experiments show that CogKR substantially outperforms previous state-of-the-art models on one-shot KG reasoning benchmarks, with relative improvements of 24.3%-29.7% on MRR. The source code is available at https://github.com/THUDM/CogKR.


Inspiration, Indeed! Alteryx Keynotes Extol the Power of *You*

#artificialintelligence

Deep down inside, you know your worth! You recognize that there's only one of you on this big blue marble, and there's no one exactly like you. Oh, maybe you get a little down sometimes, worried about the world around us; but that DNA is yours alone; and it's special. You can, and will, succeed in the brave new world of machine learning and artificial intelligence. You'll solve challenges that are interesting for you, and valuable for those around you.


5 Key Learnings To Set-up A High Impact AI Strategy

#artificialintelligence

In the following, I share the key learnings of the webinar. AI is not a secret sauce and requires lots of good data to create real value. Companies need to first separate the hype from the actual capabilities of AI, defining what AI means for them and how it might create value. Moving an entire company towards the adoption of AI is a challenging task and needs lots of educational effort. AI is not the solution to all problems. Building products do not start with thinking about AI but finding a meaningful problem that once solved adds value for the customer or user.


AI is worse at identifying household items from lower-income countries

#artificialintelligence

Object recognition algorithms sold by tech companies, including Google, Microsoft, and Amazon, perform worse when asked to identify items from lower-income countries. These are the findings of a new study conducted by Facebook's AI lab, which shows that AI bias can not only reproduce inequalities within countries, but also between them. In the study (which we spotted via Jack Clark's Import AI newsletter), researchers tested five popular off-the-shelf object recognition algorithms -- Microsoft Azure, Clarifai, Google Cloud Vision, Amazon Rekognition, and IBM Watson -- to see how well each program identified household items collected from a global dataset. The dataset included 117 categories (everything from shoes to soap to sofas) and a diverse array of household incomes and geographic locations (from a family in Burundi making $27 a month to a family in Ukraine with a monthly income of $10,090). The researchers found that the object recognition algorithms made around 10 percent more errors when asked to identify items from a household with a $50 monthly income compared to those from a household making more than $3,500.


Learning Curves for Deep Neural Networks: A Gaussian Field Theory Perspective

arXiv.org Machine Learning

A series of recent works suggest that deep neural networks (DNNs), of fixed depth, are equivalent to certain Gaussian Processes (NNGP/NTK) in the highly over-parameterized regime (width or number-of-channels going to infinity). Other works suggest that this limit is relevant for real-world DNNs. These results invite further study into the generalization properties of Gaussian Processes of the NNGP and NTK type. Here we make several contributions along this line. First, we develop a formalism, based on field theory tools, for calculating learning curves perturbatively in one over the dataset size. For the case of NNGPs, this formalism naturally extends to finite width corrections. Second, in cases where one can diagonalize the covariance-function of the NNGP/NTK, we provide analytic expressions for the asymptotic learning curves of any given target function. These go beyond the standard equivalence kernel results. Last, we provide closed analytic expressions for the eigenvalues of NNGP/NTK kernels of depth 2 fully-connected ReLU networks. For datasets on the hypersphere, the eigenfunctions of such kernels, at any depth, are hyperspherical harmonics. A simple coherent picture emerges wherein fully-connected DNNs have a strong entropic bias towards functions which are low order polynomials of the input.


Is Deep Learning an RG Flow?

arXiv.org Machine Learning

Although there has been a rapid development of practical applications, theoretical explanations of deep learning are in their infancy. A possible starting point suggests that deep learning performs a sophisticated coarse graining. Coarse graining is the foundation of the renormalization group (RG), which provides a systematic construction of the theory of large scales starting from an underlying microscopic theory. In this way RG can be interpreted as providing a mechanism to explain the emergence of large scale structure, which is directly relevant to deep learning. We pursue the possibility that RG may provide a useful framework within which to pursue a theoretical explanation of deep learning. A statistical mechanics model for a magnet, the Ising model, is used to train an unsupervised RBM. The patterns generated by the trained RBM are compared to the configurations generated through a RG treatment of the Ising model. We argue that correlation functions between hidden and visible neurons are capable of diagnosing RG-like coarse graining. Numerical experiments show the presence of RG-like patterns in correlators computed using the trained RBMs. The observables we consider are also able to exhibit important differences between RG and deep learning.


Beyond DQN/A3C: A Survey in Advanced Reinforcement Learning

#artificialintelligence

One of my favorite things about deep reinforcement learning is that, unlike supervised learning, it really, really doesn't want to work. Throwing a neural net at a computer vision problem might get you 80% of the way there. Throwing a neural net at an RL problem will probably blow something up in front of your face -- and it will blow up in a different way each time you try. A lot of the biggest challenges in RL revolve around two questions: how we interact with the environment effectively (e.g. In this post, I want to explore a few recent directions in deep RL research that attempt to address these challenges, and do so with particularly elegant parallels to human cognition. This post will begin with a quick review of two canonical deep RL algorithms -- DQN and A3C -- to provide us some intuitions to refer back to, and then jump into a deep dive on a few recent papers and breakthroughs in the categories described above.


Great Learning Expands to Europe, Asia Pacific, Africa and the Middle East

#artificialintelligence

Great Learning, India's leading Ed-tech platform for working professionals today announced that it is expanding its geographic footprint globally to Europe, Asia Pacific, Africa and the Middle East. The company will offer three of its most popular programs in Data Science & Business Analytics (PGP-DSBA - a special international variant of its business analytics program PGP-BABI ranked #1 in India for the last 4 years), Artificial Intelligence & Machine Learning (PGP-AIML) and Cyber Security (SACSP - Stanford Advanced Computer Security Program) in these geographies. Offered in association with two of the top universities of the world, Stanford University and The University of Texas, Austin, these online programs have already attracted learners from 17 countries including the UK, Singapore, South Africa, UAE, etc. These programs, designed and developed by the top-notch faculty of UT Austin and Stanford, are delivered online by Great Learning, utilizing its unique mentored-learning model where personalized mentorship is provided by expert instructors from Great Learning's 750 Global Guru network. The mentors include industry veterans from companies like Google, Microsoft, Amazon and Walmart.


Large Scale Structure of Neural Network Loss Landscapes

arXiv.org Machine Learning

There are many surprising and perhaps counter-intuitive properties of optimization of deep neural networks. We propose and experimentally verify a unified phenomenological model of the loss landscape that incorporates many of them. High dimensionality plays a key role in our model. Our core idea is to model the loss landscape as a set of high dimensional \emph{wedges} that together form a large-scale, inter-connected structure and towards which optimization is drawn. We first show that hyperparameter choices such as learning rate, network width and $L_2$ regularization, affect the path optimizer takes through the landscape in a similar ways, influencing the large scale curvature of the regions the optimizer explores. Finally, we predict and demonstrate new counter-intuitive properties of the loss-landscape. We show an existence of low loss subspaces connecting a set (not only a pair) of solutions, and verify it experimentally. Finally, we analyze recently popular ensembling techniques for deep networks in the light of our model.