Goto

Collaborating Authors

 Deep Learning


Random feature neural networks learn Black-Scholes type PDEs without curse of dimensionality

arXiv.org Machine Learning

This article investigates the use of random feature neural networks for learning Kolmogorov partial (integro-)differential equations associated to Black-Scholes and more general exponential L\'evy models. Random feature neural networks are single-hidden-layer feedforward neural networks in which only the output weights are trainable. This makes training particularly simple, but (a priori) reduces expressivity. Interestingly, this is not the case for Black-Scholes type PDEs, as we show here. We derive bounds for the prediction error of random neural networks for learning sufficiently non-degenerate Black-Scholes type models. A full error analysis is provided and it is shown that the derived bounds do not suffer from the curse of dimensionality. We also investigate an application of these results to basket options and validate the bounds numerically. These results prove that neural networks are able to \textit{learn} solutions to Black-Scholes type PDEs without the curse of dimensionality. In addition, this provides an example of a relevant learning problem in which random feature neural networks are provably efficient.


The Flip Side of the Reweighted Coin: Duality of Adaptive Dropout and Regularization

arXiv.org Machine Learning

Among the most successful methods for sparsifying deep (neural) networks are those that adaptively mask the network weights throughout training. By examining this masking, or dropout, in the linear case, we uncover a duality between such adaptive methods and regularization through the so-called "$\eta$-trick" that casts both as iteratively reweighted optimizations. We show that any dropout strategy that adapts to the weights in a monotonic way corresponds to an effective subquadratic regularization penalty, and therefore leads to sparse solutions. We obtain the effective penalties for several popular sparsification strategies, which are remarkably similar to classical penalties commonly used in sparse optimization. Considering variational dropout as a case study, we demonstrate similar empirical behavior between the adaptive dropout method and classical methods on the task of deep network sparsification, validating our theory.


PopSkipJump: Decision-Based Attack for Probabilistic Classifiers

arXiv.org Machine Learning

Most current classifiers are vulnerable to adversarial examples, small input perturbations that change the classification output. Many existing attack algorithms cover various settings, from white-box to black-box classifiers, but typically assume that the answers are deterministic and often fail when they are not. We therefore propose a new adversarial decision-based attack specifically designed for classifiers with probabilistic outputs. It is based on the HopSkipJump attack by Chen et al. (2019, arXiv:1904.02144v5 ), a strong and query efficient decision-based attack originally designed for deterministic classifiers. Our P(robabilisticH)opSkipJump attack adapts its amount of queries to maintain HopSkipJump's original output quality across various noise levels, while converging to its query efficiency as the noise level decreases. We test our attack on various noise models, including state-of-the-art off-the-shelf randomized defenses, and show that they offer almost no extra robustness to decision-based attacks. Code is available at https://github.com/cjsg/PopSkipJump .


Convergence and Alignment of Gradient Descent with Random Back Propagation Weights

arXiv.org Machine Learning

Stochastic gradient descent with backpropagation is the workhorse of artificial neural networks. It has long been recognized that backpropagation fails to be a biologically plausible algorithm. Fundamentally, it is a non-local procedure -- updating one neuron's synaptic weights requires knowledge of synaptic weights or receptive fields of downstream neurons. This limits the use of artificial neural networks as a tool for understanding the biological principles of information processing in the brain. Lillicrap et al. (2016) propose a more biologically plausible "feedback alignment" algorithm that uses random and fixed backpropagation weights, and show promising simulations. In this paper we study the mathematical properties of the feedback alignment procedure by analyzing convergence and alignment for two-layer networks under squared error loss. In the overparameterized setting, we prove that the error converges to zero exponentially fast, and also that regularization is necessary in order for the parameters to become aligned with the random backpropagation weights. Simulations are given that are consistent with this analysis and suggest further generalizations. These results contribute to our understanding of how biologically plausible algorithms might carry out weight learning in a manner different from Hebbian learning, with performance that is comparable with the full non-local backpropagation algorithm.


Cybersecurity experts face a new challenge: AI capable of tricking them

#artificialintelligence

If you use such social media websites as Facebook and Twitter, you may have come across posts flagged with warnings about misinformation. So far, most misinformation – flagged and unflagged – has been aimed at the general public. Imagine the possibility of misinformation – information that is false or misleading – in scientific and technical fields like cybersecurity, public safety and medicine. There is growing concern about misinformation spreading in these critical fields as a result of common biases and practices in publishing scientific literature, even in peer-reviewed research papers. As a graduate student and as faculty members doing research in cybersecurity, we studied a new avenue of misinformation in the scientific community.


Validate computer vision deep learning models

#artificialintelligence

This code pattern is part of the Getting started with IBM Maximo Visual Inspection learning path. After a deep learning computer vision model is trained and deployed, it is often necessary to periodically (or continuously) evaluate the model with new test data. This developer code pattern provides a Jupyter Notebook that will take test images with known "ground-truth" categories and evaluate the inference results versus the truth. We will use a Jupyter Notebook to evaluate an IBM Maximo Visual Inspection image classification model. You can train a model using the provided example or test your own deployed model.


Machine Learning Basics Everyone Should Know - InformationWeek

#artificialintelligence

AI is seeping into just about everything, from consumer products to industrial equipment. As enterprises utilize AI to become more competitive, more of them are taking advantage of machine learning to accomplish more in less time, reduce costs and discover something whether a drug or a latent market desire. While there's no need for non-data scientists to understand how machine learning (ML) works, they should understand enough to use basic terminology correctly. Although the scope of ML extends considerably past what's possible to cover in this short article, following are some of the fundamentals. Before one can grasp machine learning concepts, they need to understand what machine learning terms mean.


Probabilistic Deep Learning with TensorFlow 2

#artificialintelligence

About this Course 42,788 recent views Welcome to this course on Probabilistic Deep Learning with TensorFlow! This course builds on the foundational concepts and skills for TensorFlow taught in the first two courses in this specialisation, and focuses on the probabilistic approach to deep learning. This is an increasingly important area of deep learning that aims to quantify the noise and uncertainty that is often present in real world datasets. This is a crucial aspect when using deep learning models in applications such as autonomous vehicles or medical diagnoses; we need the model to know what it doesn't know. You will learn how to develop probabilistic models with TensorFlow, making particular use of the TensorFlow Probability library, which is designed to make it easy to combine probabilistic models with deep learning.


Keeping a closer eye on seabirds with drones and artificial intelligence

#artificialintelligence

Using drones and artificial intelligence to monitor large colonies of seabirds can be as effective as traditional on-the-ground methods, while reducing costs, labor and the risk of human error, a new study finds. Scientists at Duke University and the Wildlife Conservation Society (WCS) used a deep-learning algorithm--a form of artificial intelligence--to analyze more than 10,000 drone images of mixed colonies of seabirds in the Falkland Islands off Argentina's coast. The Falklands, also known as the Malvinas, are home to the world's largest colonies of black-browed albatrosses (Thalassarche melanophris) and second-largest colonies of southern rockhopper penguins (Eudyptes c. chrysocome). Hundreds of thousands of birds breed on the islands in densely interspersed groups. The deep-learning algorithm correctly identified and counted the albatrosses with 97% accuracy and the penguins with 87%.


Engineers Apply Physics-informed Machine Learning To Solar Cell Production - AI Summary

#artificialintelligence

Despite the recent advances in the power conversion efficiency of organic solar cells, insights into the processing-driven thermo-mechanical stability of bulk heterojunction active layers are helping to advance the field. Lehigh University engineer Ganesh Balasubramanian, like many others, wondered if there were ways to improve the design of solar cells to make them more efficient? Balasubramanian, an associate professor of Mechanical Engineering and Mechanics, studies the basic physics of the materials at the heart of solar energy conversion – the organic polymers passing electrons from molecule to molecule so they can be stored and harnessed – as well as the manufacturing processes that produce commercial solar cells. Using the Frontera supercomputer at the Texas Advanced Computing Center (TACC) – one of the most powerful on the planet – Balasubramanian and his graduate student Joydeep Munshi have been running molecular models of organic solar cell production processes, and designing a framework to determine the optimal engineering choices. "When engineers make solar cells, they mix two organic molecules in a solvent and evaporate the solvent to create a mixture which helps with the exciton conversion and electron transport," Balasubramanian said.