Goto

Collaborating Authors

 Deep Learning


PatchUp: A Regularization Technique for Convolutional Neural Networks

arXiv.org Machine Learning

Large capacity deep learning models are often prone to a high generalization gap when trained with a limited amount of labeled training data. A recent class of methods to address this problem uses various ways to construct a new training sample by mixing a pair (or more) of training samples. We propose PatchUp, a hidden state block-level regularization technique for Convolutional Neural Networks (CNNs), that is applied on selected contiguous blocks of feature maps from a random pair of samples. Our approach improves the robustness of CNN models against the manifold intrusion problem that may occur in other state-of-the-art mixing approaches like Mixup and CutMix. Moreover, since we are mixing the contiguous block of features in the hidden space, which has more dimensions than the input space, we obtain more diverse samples for training towards different dimensions. Our experiments on CIFAR-10, CIFAR-100, and SVHN datasets with PreactResnet18, PreactResnet34, and WideResnet-28-10 models show that PatchUp improves upon, or equals, the performance of current state-of-the-art regularizers for CNNs. We also show that PatchUp can provide better generalization to affine transformations of samples and is more robust against adversarial attacks.


A Baseline for Shapley Values in MLPs: from Missingness to Neutrality

arXiv.org Machine Learning

Being able to explain a prediction as well as having a model that performs well are paramount in many machine learning applications. Deep neural networks have gained momentum recently on the basis of their accuracy, however these are often criticised to be black-boxes. Many authors have focused on proposing methods to explain their predictions. Among these explainability methods, feature attribution methods have been favoured for their strong theoretical foundation: the Shapley value. A limitation of Shapley value is the need to define a baseline (aka reference point) representing the missingness of a feature. In this paper, we present a method to choose a baseline based on a neutrality value: a parameter defined by decision makers at which their choices are determined by the returned value of the model being either below or above it. Based on this concept, we theoretically justify these neutral baselines and find a way to identify them for MLPs. Then, we experimentally demonstrate that for a binary classification task, using a synthetic dataset and a dataset coming from the financial domain, the proposed baselines outperform, in terms of local explanability power, standard ways of choosing them.


Pruning via Iterative Ranking of Sensitivity Statistics

arXiv.org Machine Learning

With the introduction of SNIP [arXiv:1810.02340v2], it has been demonstrated that modern neural networks can effectively be pruned before training. Yet, its sensitivity criterion has since been criticized for not propagating training signal properly or even disconnecting layers. As a remedy, GraSP [arXiv:2002.07376v1] was introduced, compromising on simplicity. However, in this work we show that by applying the sensitivity criterion iteratively in smaller steps - still before training - we can improve its performance without difficult implementation. As such, we introduce 'SNIP-it'. We then demonstrate how it can be applied for both structured and unstructured pruning, before and/or during training, therewith achieving state-of-the-art sparsity-performance trade-offs. That is, while already providing the computational benefits of pruning in the training process from the start. Furthermore, we evaluate our methods on robustness to overfitting, disconnection and adversarial attacks as well.


Model Linkage Selection for Cooperative Learning

arXiv.org Machine Learning

Rapid developments in data collecting devices and computation platforms produce an emerging number of learners and data modalities in many scientific domains. We consider the setting in which each learner holds a pair of parametric statistical model and a specific data source, with the goal of integrating information across a set of learners to enhance the prediction accuracy of a specific learner. One natural way to integrate information is to build a joint model across a set of learners that shares common parameters of interest. However, the parameter sharing patterns across a set of learners are not known a priori. Misspecifying the parameter sharing patterns and the parametric statistical model for each learner yields a biased estimator and degrades the prediction accuracy of the joint model. In this paper, we propose a novel framework for integrating information across a set of learners that is robust against model misspecification and misspecified parameter sharing patterns. The main crux is to sequentially incorporates additional learners that can enhance the prediction accuracy of an existing joint model based on a user-specified parameter sharing patterns across a set of learners, starting from a model with one learner. Theoretically, we show that the proposed method can data-adaptively select the correct parameter sharing patterns based on a user-specified parameter sharing patterns, and thus enhances the prediction accuracy of a learner. Extensive numerical studies are performed to evaluate the performance of the proposed method.


Machine learning based digital twin for dynamical systems with multiple time-scales

arXiv.org Machine Learning

Digital twin technology has a huge potential for widespread applications in different industrial sectors such as infrastructure, aerospace, and automotive. However, practical adoptions of this technology have been slower, mainly due to a lack of application-specific details. Here we focus on a digital twin framework for linear single-degree-of-freedom structural dynamic systems evolving in two different operational time scales in addition to its intrinsic dynamic time-scale. Our approach strategically separates into two components -- (a) a physics-based nominal model for data processing and response predictions, and (b) a data-driven machine learning model for the time-evolution of the system parameters. The physics-based nominal model is system-specific and selected based on the problem under consideration. On the other hand, the data-driven machine learning model is generic. For tracking the multi-scale evolution of the system parameters, we propose to exploit a mixture of experts as the data-driven model. Within the mixture of experts model, Gaussian Process (GP) is used as the expert model. The primary idea is to let each expert track the evolution of the system parameters at a single time-scale. For learning the hyperparameters of the `mixture of experts using GP', an efficient framework the exploits expectation-maximization and sequential Monte Carlo sampler is used. Performance of the digital twin is illustrated on a multi-timescale dynamical system with stiffness and/or mass variations. The digital twin is found to be robust and yields reasonably accurate results. One exciting feature of the proposed digital twin is its capability to provide reasonable predictions at future time-steps. Aspects related to the data quality and data quantity are also investigated.


Facial Recognition Bans: What Do They Mean For AI (Artificial Intelligence)?

#artificialintelligence

This week IBM, Microsoft and Amazon announced that they would suspend the sale of their facial recognition technology to law enforcement agencies. But the moves from the tech giants also illustrate the inherent risks of AI, especially when it comes to bias and the potential for invasion of privacy. Note that there are already indications that Congress will take action to regulate the technology. In the meantime, many cities have already instituted bans, such San Francisco. Because of the advances of deep learning and faster systems for processing enormous amounts of data, facial recognition has certainly seen major strides over the past decade.


Loss landscapes and the blessing of dimensionality

#artificialintelligence

"Life requires Movement" -- (Aristotle, 4th century BC) The more life there is, the more flexibility there is. The more fluid you are, the more you are alive." "Life is movement, movement is change" -- beginning of a quote by Neale Donald Walsch "If life boils down to one thing, it's movement. To live is to keep moving" -- Jerry Seinfeld As many well known people have told us throughout history, life is movement. Life is changing your state in a proactive way, going from A to B. Life is also an expensive process and that makes the process of going from A to B a delicate, fascinating process that needs to be optimized. And that's where we begin this article. And although we will focus throughout the following sections on deep learning and A.I, we will be touching simultaneously on universal themes and principles that go to the core of what means being alive. In a process that depends on a very large number of parameters, which makes it multidimensional. Which makes it hard to visualize for beings that operate in only 3 dimensions (4 with time). Which is the whole point of this article. So let's begin this ride where it all begins, with movement. Say we want to go from A to B in regards to some objective. Some of these challenges will take minutes, other hours, others days and some of them years. Some of them depend on a moderate number of factors, others depend on a massive number of them. We want to optimize these and infinite other challenges and the objective is always going from A to B. You may also combine many challenges and see life itself as a massive fractal made of optimization processes at different scales. Going from A to B could be tackled in different ways. We could do it very systematically, trying lots of possibilities. Or we could try to find the most efficient way to get there as soon as possible.


Implementation of model explainability for a basic brain tumor detection using convolutional neural networks on MRI slices

#artificialintelligence

While neural networks gain popularity in medical research, attempts to make the decisions of a model explainable are often only made towards the end of the development process once a high predictive accuracy has been achieved. In order to assess the advantages of implementing features to increase explainability early in the development process, we trained a neural network to differentiate between MRI slices containing either a vestibular schwannoma, a glioblastoma, or no tumor. Making the decisions of a network more explainable helped to identify potential bias and choose appropriate training data. Model explainability should be considered in early stages of training a neural network for medical purposes as it may save time in the long run and will ultimately help physicians integrate the network's predictions into a clinical decision.


A 2020 Guide To Text Moderation with NLP and Deep Learning

#artificialintelligence

In this article, we will look at toxic speech detection, the problem of text moderation and understand the different challenges that one might encounter trying to automate the process. We look at several NLP and deep learning approaches to solve the problem and finally implement a toxic speech classifier using BERT embeddings. As of June 2019 there are now over 4.4 billion internet users. According to the latest Domo Data Never Sleeps report, Twitter users send 511,200 tweets per minute. While that happens, TikTok gets banned in Indonesia, Discord sees an increasing number of neo-Nazi posts, tech and film celebrity accounts get hacked so hackers can spurt out several racist slurs and hate speech volumes rise in India on facebook due to the controversial Citizenship Amendment Act (CAA). Social media continues to be used by several to incite violence, spread hate and target minorities based on religion, sex, race and disabilities.


Artificial Intelligence, Machine Learning, and Deep Learning -- What the Difference?

#artificialintelligence

Electricity has changed how the world operated. It changes transportation, manufacturing, agriculture, and even health care. For example, before the invention of electric lighting, humans were limited to daytime activities, because at night it was dark, only people who could afford gas lamps could do activities. Compared to now, we can still do activities at night because it is illuminated by electric lights. AI is expected to have a similar effect.