Goto

Collaborating Authors

 Deep Learning


Compressing Deep Neural Networks via Layer Fusion

arXiv.org Machine Learning

This paper proposes \textit{layer fusion} - a model compression technique that discovers which weights to combine and then fuses weights of similar fully-connected, convolutional and attention layers. Layer fusion can significantly reduce the number of layers of the original network with little additional computation overhead, while maintaining competitive performance. From experiments on CIFAR-10, we find that various deep convolution neural networks can remain within 2\% accuracy points of the original networks up to a compression ratio of 3.33 when iteratively retrained with layer fusion. For experiments on the WikiText-2 language modelling dataset where pretrained transformer models are used, we achieve compression that leads to a network that is 20\% of its original size while being within 5 perplexity points of the original network. We also find that other well-established compression techniques can achieve competitive performance when compared to their original networks given a sufficient number of retraining steps. Generally, we observe a clear inflection point in performance as the amount of compression increases, suggesting a bound on the amount of compression that can be achieved before an exponential degradation in performance.


Approaches to Fraud Detection on Credit Card Transactions Using Artificial Intelligence Methods

arXiv.org Machine Learning

Credit card fraud is an ongoing problem for almost all industries in the world, and it raises millions of dollars to the global economy each year. Therefore, there is a number of research either completed or proceeding in order to detect these kinds of frauds in the industry. These researches generally use rule-based or novel artificial intelligence approaches to find eligible solutions. The ultimate goal of this paper is to summarize state-of-the-art approaches to fraud detection using artificial intelligence and machine learning techniques. While summarizing, we will categorize the common problems such as imbalanced dataset, real time working scenarios, and feature engineering challenges that almost all research works encounter, and identify general approaches to solve them. The imbalanced dataset problem occurs because the number of legitimate transactions is much higher than the fraudulent ones whereas applying the right feature engineering is substantial as the features obtained from the industries are limited, and applying feature engineering methods and reforming the dataset is crucial. Also, adapting the detection system to real time scenarios is a challenge since the number of credit card transactions in a limited time period is very high. In addition, we will discuss how evaluation metrics and machine learning methods differentiate among each research. NTRODUCTION The number of cashless transactions is at its peak point since the beginning of the digital era and it is most likely to increase in the future.


Clarinet: A One-step Approach Towards Budget-friendly Unsupervised Domain Adaptation

arXiv.org Machine Learning

In unsupervised domain adaptation (UDA), classifiers for the target domain are trained with massive true-label data from the source domain and unlabeled data from the target domain. However, it may be difficult to collect fully-true-label data in a source domain given a limited budget. To mitigate this problem, we consider a novel problem setting where the classifier for the target domain has to be trained with complementary-label data from the source domain and unlabeled data from the target domain named budget-friendly UDA (BFUDA). The key benefit is that it is much less costly to collect complementary-label source data (required by BFUDA) than collecting the true-label source data (required by ordinary UDA). To this end, the complementary label adversarial network (CLARINET) is proposed to solve the BFUDA problem. CLARINET maintains two deep networks simultaneously, where one focuses on classifying complementary-label source data and the other takes care of the source-to-target distributional adaptation. Experiments show that CLARINET significantly outperforms a series of competent baselines.


A regularized deep matrix factorized model of matrix completion for image restoration

arXiv.org Machine Learning

It has been an important approach of using matrix completion to perform image restoration. Most previous works on matrix completion focus on the low-rank property by imposing explicit constraints on the recovered matrix, such as the constraint of the nuclear norm or limiting the dimension of the matrix factorization component. Recently, theoretical works suggest that deep linear neural network has an implicit bias towards low rank on matrix completion. However, low rank is not adequate to reflect the intrinsic characteristics of a natural image. Thus, algorithms with only the constraint of low rank are insufficient to perform image restoration well. In this work, we propose a Regularized Deep Matrix Factorized (RDMF) model for image restoration, which utilizes the implicit bias of the low rank of deep neural networks and the explicit bias of total variation. We demonstrate the effectiveness of our RDMF model with extensive experiments, in which our method surpasses the state of art models in common examples, especially for the restoration from very few observations. Our work sheds light on a more general framework for solving other inverse problems by combining the implicit bias of deep learning with explicit regularization.


Incremental Training of Graph Neural Networks on Temporal Graphs under Distribution Shift

arXiv.org Machine Learning

Current graph neural networks (GNNs) are promising, especially when the entire graph is known for training. However, it is not yet clear how to efficiently train GNNs on temporal graphs, where new vertices, edges, and even classes appear over time. We face two challenges: First, shifts in the label distribution (including the appearance of new labels), which require adapting the model. Second, the growth of the graph, which makes it, at some point, infeasible to train over all vertices and edges. We address these issues by applying a sliding window technique, i.e., we incrementally train GNNs on limited window sizes and analyze their performance. For our experiments, we have compiled three new temporal graph datasets based on scientific publications and evaluate isotropic and anisotropic GNN architectures. Our results show that both GNN types provide good results even for a window size of just 1 time step. With window sizes of 3 to 4 time steps, GNNs achieve at least 95% accuracy compared to using the entire timeline of the graph. With window sizes of 6 or 8, at least 99% accuracy could be retained. These discoveries have direct consequences for training GNNs over temporal graphs. We provide the code (https://github.com/Incremental-GNNs) and the newly compiled datasets (https://zenodo.org/record/3764770) for reproducibility and reuse.


Leveraging the Power of AI for European Cultural Heritage

#artificialintelligence

BARCELONA, Spain, July 28, 2020 -- AI will now be well-versed in cultural heritage due to a new EU-funded project called Saint George on a Bike. Composed of researchers from the Barcelona Supercomputing Center (BSC) and Europeana Foundation, the project has begun training natural language processing and deep learning algorithms in culture, symbols, and historical context with the aim of automatically generating rich metadata for hundreds of thousands of images from various European cultural heritage repositories. Training AI to be aware of cultural heritage contexts is not as simple as teaching it to identify different objects in a picture. Saint George on a Bike is fine-tuning the algorithms so that it "thinks" in context and according to time parameters. "The AI we are developing will be able to tell whether a painting shows Saint George on a horse or a bike," said Maria Cristina Marinescu, senior researcher at BSC and coordinator of the Saint George on a Bike project. "This is not as easy as it sounds because the shapes are similar.


Looking into the black box of deep learning - ScienceBlog.com

#artificialintelligence

Deep learning systems are revolutionizing technology around us, from voice recognition that pairs you with your phone to autonomous vehicles that are increasingly able to see and recognize obstacles ahead. But much of this success involves trial and error when it comes to the deep learning networks themselves. A group of MIT researchers recently reviewed their contributions to a better theoretical understanding of deep learning networks, providing direction for the field moving forward. "Deep learning was in some ways an accidental discovery," explains Tommy Poggio, investigator at the McGovern Institute for Brain Research, director of the Center for Brains, Minds, and Machines (CBMM), and the Eugene McDermott Professor in Brain and Cognitive Sciences. "We still do not understand why it works. A theoretical framework is taking form, and I believe that we are now close to a satisfactory theory. It is time to stand back and review recent insights."


Exxact Extends Deep Learning Infrastructure Solutions with NVIDIA DGX A100 Systems

#artificialintelligence

The NVIDIA DGX A100 is a high-performance computing system for AI training, inference and analytics. It sets a new bar for compute density, packing 5 petaFLOPS of AI performance into a 6U form factor, replacing legacy infrastructure silos with one flexible platform that can support every AI workload. "With the NVIDIA DGX A100, NVIDIA has really changed the game for AI in terms of extreme performance, scale and flexibility. By offering colocation services and flexible lease options, we're making this technology more accessible than ever before," said Jason Chen, Vice President of Exxact Corporation. More than just a server, the DGX A100 integrates exclusive access to the Exxact team of AI-fluent experts that offer prescriptive planning, deployment, and optimization expertise to help fast-track AI transformation. Available now, the NVIDIA DGX A100 can be bundled with an optional three-year warranty and support package to improve productivity by reducing downtime on production systems.


Commentary: Optimizing a truck fleet using artificial intelligence - FreightWaves

#artificialintelligence

The views expressed here are solely those of the author and do not necessarily represent the views of FreightWaves or its affiliates. Author's Disclosure: I am not an investor in Optimal Dynamics, either personally or through REFASHIOND Ventures. I have no other financial relationship with Optimal Dynamics. On July 7 I started a series on AI in Supply Chain (#AIinSupplyChain). The first article in the series profiled Optimal Dynamics, a startup that has launched a product to automatically optimize operations for large trucking fleets.


Uber ATG Open-Sources Neuropod DL Inference Engine

#artificialintelligence

Every Neuropod model implements a problem definition -- a formal description of a problem for models to solve. As a result, any models that solve the same problem are interchangeable even if they use different frameworks. Existing models can be wrapped in a Neuropod package, which contains the original model along with metadata, test data, and custom ops if any. Since its internal release in early 2019, hundred of Neuropod models have been deployed across Uber ATG, Uber AI, and the core Uber business -- including models for demand forecasting, estimated time of arrival (ETA) prediction for rides, menu transcription for Uber Eats, and object detection models for self-driving vehicles. Neuropod makes it easy for researchers to build models in a framework of their choosing while also simplifying product-ionization of these models, says the company.