Goto

Collaborating Authors

 Deep Learning


SHINE: SHaring the INverse Estimate from the forward pass for bi-level optimization and implicit models

arXiv.org Machine Learning

In recent years, implicit deep learning has emerged as a method to increase the depth of deep neural networks. While their training is memory-efficient, they are still significantly slower to train than their explicit counterparts. In Deep Equilibrium Models (DEQs), the training is performed as a bi-level problem, and its computational complexity is partially driven by the iterative inversion of a huge Jacobian matrix. In this paper, we propose a novel strategy to tackle this computational bottleneck from which many bi-level problems suffer. The main idea is to use the quasi-Newton matrices from the forward pass to efficiently approximate the inverse Jacobian matrix in the direction needed for the gradient computation. We provide a theorem that motivates using our method with the original forward algorithms. In addition, by modifying these forward algorithms, we further provide theoretical guarantees that our method asymptotically estimates the true implicit gradient. We empirically study this approach in many settings, ranging from hyperparameter optimization to large Multiscale DEQs applied to CIFAR and ImageNet. We show that it reduces the computational cost of the backward pass by up to two orders of magnitude. All this is achieved while retaining the excellent performance of the original models in hyperparameter optimization and on CIFAR, and giving encouraging and competitive results on ImageNet.


Locally Valid and Discriminative Confidence Intervals for Deep Learning Models

arXiv.org Machine Learning

Crucial for building trust in deep learning models for critical real-world applications is efficient and theoretically sound uncertainty quantification, a task that continues to be challenging. Useful uncertainty information is expected to have two key properties: It should be valid (guaranteeing coverage) and discriminative (more uncertain when the expected risk is high). Moreover, when combined with deep learning (DL) methods, it should be scalable and affect the DL model performance minimally. Most existing Bayesian methods lack frequentist coverage guarantees and usually affect model performance. The few available frequentist methods are rarely discriminative and/or violate coverage guarantees due to unrealistic assumptions. Moreover, many methods are expensive or require substantial modifications to the base neural network. Building upon recent advances in conformal prediction and leveraging the classical idea of kernel regression, we propose Locally Valid and Discriminative confidence intervals (LVD), a simple, efficient and lightweight method to construct discriminative confidence intervals (CIs) for almost any DL model. With no assumptions on the data distribution, such CIs also offer finite-sample local coverage guarantees (contrasted to the simpler marginal coverage). Using a diverse set of datasets, we empirically verify that besides being the only locally valid method, LVD also exceeds or matches the performance (including coverage rate and prediction accuracy) of existing uncertainty quantification methods, while offering additional benefits in scalability and flexibility.


Neural Network Training Using $\ell_1$-Regularization and Bi-fidelity Data

arXiv.org Machine Learning

With the capability of accurately representing a functional relationship between the inputs of a physical system's model and output quantities of interest, neural networks have become popular for surrogate modeling in scientific applications. However, as these networks are over-parameterized, their training often requires a large amount of data. To prevent overfitting and improve generalization error, regularization based on, e.g., $\ell_1$- and $\ell_2$-norms of the parameters is applied. Similarly, multiple connections of the network may be pruned to increase sparsity in the network parameters. In this paper, we explore the effects of sparsity promoting $\ell_1$-regularization on training neural networks when only a small training dataset from a high-fidelity model is available. As opposed to standard $\ell_1$-regularization that is known to be inadequate, we consider two variants of $\ell_1$-regularization informed by the parameters of an identical network trained using data from lower-fidelity models of the problem at hand. These bi-fidelity strategies are generalizations of transfer learning of neural networks that uses the parameters learned from a large low-fidelity dataset to efficiently train networks for a small high-fidelity dataset. We also compare the bi-fidelity strategies with two $\ell_1$-regularization methods that only use the high-fidelity dataset. Three numerical examples for propagating uncertainty through physical systems are used to show that the proposed bi-fidelity $\ell_1$-regularization strategies produce errors that are one order of magnitude smaller than those of networks trained only using datasets from the high-fidelity models.



Under the Hood of Modern Machine and Deep Learning

#artificialintelligence

In this chapter, we investigate whether unique, optimal decision boundaries can be found. In order to do so, we first have to revisit several fundamental mathematical principles. Regularization is a mathematical tool, which allows us to find unique solutions even for highly ill-posed problems. In order to use this trick, we review norms and how they can be used to steer regression problems. Rosenblatt's Perceptron and Multi-Layer Perceptrons which are also called Artificial Neural Networks inherently suffer from this ill-posedness.


This mathematical brain model may pave the way for more human-like AI

#artificialintelligence

Last week, Google Research held an online workshop on the conceptual understanding of deep learning. The workshop, which featured presentations by award-winning computer scientists and neuroscientists, discussed how new findings in deep learning and neuroscience can help create better artificial intelligence systems. While all the presentations and discussions were worth watching (and I might revisit them again in the coming weeks), one, in particular, stood out for me: A talk on word representations in the brain by Christos Papadimitriou, professor of computer science at the University of Columbia. In his presentation, Papadimitriou, a recipient of the Gรถdel Prize and Knuth Prize, discussed how our growing understanding of information-processing mechanisms in the brain might help create algorithms that are more robust in understanding and engaging in conversations. Papadimitriou presented a simple and efficient model that explains how different areas of the brain inter-communicate to solve cognitive problems.


Artificial Intelligence Identifies Netflix And Microsoft Among Today's Trending Stocks

#artificialintelligence

Let's see why these giants of industry (and pandemic profits) are circulating in investors' portfolios. Q.ai runs daily factor models to get the most up-to-date reading on stocks and ETFs. Our deep-learning algorithms use Artificial Intelligence (AI) technology to provide an in-depth, intelligence-based look at a company โ€“ so you don't have to do the digging yourself. Sign up for the free Forbes AI Investor newsletter here to join an exclusive AI investing community and get premium investing ideas before markets open. Netflix, Inc NFLX ticked up 0.2% on Wednesday to $502.36 per share, ending the day with 2.46 million trades.


This is a great moment to look for a new job in Artificial Intelligence.

#artificialintelligence

Artificial intelligence (AI) is essential because it allows the software to perform human capacities such as understanding, reasoning, planning, communication, and perception in an increasingly effective, efficient, and low-cost manner. In most business sectors, automating these skills opens up new opportunities. With the significant evolution of algorithms, AI is already a reality. Deep Learning algorithms such as Convolutional Neural Networks (CNNs), for example, have significantly improved computers' ability to recognize objects in images. In addition, Recurrent Neural Networks (RNNs) algorithms produce voice recognition systems that outperform humans.


Web Scraping Images From Google With Selenium

#artificialintelligence

Hello & Welcome everyone after a long gap. We have been off for a while now but we are back with whole lot of interesting stuff to read & learn about. In today's post we are gonna see Web Scraping Images from Google with Selenium . This blog is in continuation to our previous blog here where we used Haarcascade and opencv for mask detection. In this post i wanted to try a deep learning approach to overcome the shortcomings of the previous blog.


This is a great moment to look for a new job in Artificial Intelligence.

#artificialintelligence

Artificial intelligence (AI) is essential because it allows the software to perform human capacities such as understanding, reasoning, planning, communication, and perception in an increasingly effective, efficient, and low-cost manner. In most business sectors, automating these skills opens up new opportunities. With the significant evolution of algorithms, AI is already a reality. Deep Learning algorithms such as Convolutional Neural Networks (CNNs), for example, have significantly improved computers' ability to recognize objects in images. In addition, Recurrent Neural Networks (RNNs) algorithms produce voice recognition systems that outperform humans.