Goto

Collaborating Authors

 Deep Learning


DeepMind releases database with AI predictions for every human protein shape

#artificialintelligence

DeepMind released a free, open-source, big-deal database last week containing AI predictions for the shapes of every protein in the human body. Not only is it the most complete picture of the human proteome (full set of proteins) to date, according to the London-based AI lab--it's also "doubling humanity's accumulated knowledge of high-accuracy human protein structures." Deepening our understanding of protein structures can lead to major leaps forward in understanding diseases, as well as in drug and vaccine development. That could help allay anything from neglected diseases to the next pandemic. Recap: In December 2020, AlphaFold, DeepMind's neural network, made a breakthrough in protein folding--a biological mystery that's puzzled scientists for 50 years.


GradCAM in PyTorch

#artificialintelligence

In this article, we are going to learn how to plot GradCam [1] in PyTorch. To get the GradCam outputs, we need the activation maps and the gradients of those activation maps. Let us jump straight into the code!! We are going to use hooks to get the activation maps and the gradients from the desired layer and tensor, respectively. For this tutorial, we are going to take the activation maps from layer4 of ResNet50 and gradients with respect to the output tensor of the same.


Council Post: How An Avalanche Of Data Led To New Trends In AI Software Modernization Approaches

#artificialintelligence

Evgeniy is a specialist in software development, technological entrepreneurship and emerging technologies. In recent years, companies' growing focus on big data has led to increased digitalization demands. The avalanche of data has forced businesses to reconsider software modernization approaches. With that in mind, let's look at how enterprises use AI in intelligent analysis, hyperautomation and cybersecurity in the world of big data. Data orientation is the future of business, and the survival of companies depends on efficiently processing external and internal information.


Hit Song Prediction

#artificialintelligence

Can you predict the success of a song, just by listening to it? I know! at least we try to do it many times, and a lot of times our predictions do turn out to be true. While we do consider many things and most importantly the emotions involved. Can we expect a deep learning model to predict that!?. In this article, we will figure out how though.


DeepMind will soon publish the structure of every protein known to science

#artificialintelligence

DeepMind, a sister company of Google, is giving the world access to a massive protein structure database -- a gift that has the potential to revolutionize scientific research. "This will be one of the most important datasets since the mapping of the Human Genome," Ewan Birney, deputy director general of the European Molecular Biology Laboratory, which partnered with DeepMind on the database, said in a press release. Protein structure: Proteins are molecules that are hugely important to the functioning of living organisms, including humans -- practically everything we're made of and everything our cells do is determined by our proteins. "It's the most significant contribution AI has made to advancing scientific knowledge to date." Every protein is made up of a long string of hundreds or even thousands of chemical compounds called amino acids, and the way that ribbon folds on itself determines the protein's function.


Artificial Intelligence learns better when distracted

#artificialintelligence

Convolutional Neural Networks (CNNs) are a form of bio-inspired deep learning in artificial intelligence. The interaction of thousands of'neurons' mimics the way our brain learns to recognize images. 'These CNNs are successful, but we don't fully understand how they work', says Estefanía Talavera Martinez, lecturer and researcher at the Bernoulli Institute for Mathematics, Computer Science and Artificial Intelligence of the University of Groningen in the Netherlands. She has made use of CNNs herself to analyse images made by wearable cameras in the study of human behaviour. Among other things, Talavera Martinez has been studying our interactions with food, so she wanted the system to recognize the different settings in which people encounter food.


A purely data-driven framework for prediction, optimization, and control of networked processes: application to networked SIS epidemic model

arXiv.org Artificial Intelligence

Networks are landmarks of many complex phenomena where interweaving interactions between different agents transform simple local rule-sets into nonlinear emergent behaviors. While some recent studies unveil associations between the network structure and the underlying dynamical process, identifying stochastic nonlinear dynamical processes continues to be an outstanding problem. Here we develop a simple data-driven framework based on operator-theoretic techniques to identify and control stochastic nonlinear dynamics taking place over large-scale networks. The proposed approach requires no prior knowledge of the network structure and identifies the underlying dynamics solely using a collection of two-step snapshots of the states. This data-driven system identification is achieved by using the Koopman operator to find a low dimensional representation of the dynamical patterns that evolve linearly. Further, we use the global linear Koopman model to solve critical control problems by applying to model predictive control (MPC)--typically, a challenging proposition when applied to large networks. We show that our proposed approach tackles this by converting the original nonlinear programming into a more tractable optimization problem that is both convex and with far fewer variables.


Bilevel Optimization for Machine Learning: Algorithm Design and Convergence Analysis

arXiv.org Machine Learning

Bilevel optimization has become a powerful framework in various machine learning applications including meta-learning, hyperparameter optimization, and network architecture search. There are generally two classes of bilevel optimization formulations for machine learning: 1) problem-based bilevel optimization, whose inner-level problem is formulated as finding a minimizer of a given loss function; and 2) algorithm-based bilevel optimization, whose inner-level solution is an output of a fixed algorithm. For the first class, two popular types of gradient-based algorithms have been proposed for hypergradient estimation via approximate implicit differentiation (AID) and iterative differentiation (ITD). Algorithms for the second class include the popular model-agnostic meta-learning (MAML) and almost no inner loop (ANIL). However, the convergence rate and fundamental limitations of bilevel optimization algorithms have not been well explored. This thesis provides a comprehensive convergence rate analysis for bilevel algorithms in the aforementioned two classes. We further propose principled algorithm designs for bilevel optimization with higher efficiency and scalability. For the problem-based formulation, we provide a convergence rate analysis for AID- and ITD-based bilevel algorithms. We then develop acceleration bilevel algorithms, for which we provide shaper convergence analysis with relaxed assumptions. We also provide the first lower bounds for bilevel optimization, and establish the optimality by providing matching upper bounds under certain conditions. We finally propose new stochastic bilevel optimization algorithms with lower complexity and higher efficiency in practice. For the algorithm-based formulation, we develop a theoretical convergence for general multi-step MAML and ANIL, and characterize the impact of parameter selections and loss geometries on the their complexities.


Learning with Noisy Labels via Sparse Regularization

arXiv.org Machine Learning

Learning with noisy labels is an important and challenging task for training accurate deep neural networks. Some commonly-used loss functions, such as Cross Entropy (CE), suffer from severe overfitting to noisy labels. Robust loss functions that satisfy the symmetric condition were tailored to remedy this problem, which however encounter the underfitting effect. In this paper, we theoretically prove that \textbf{any loss can be made robust to noisy labels} by restricting the network output to the set of permutations over a fixed vector. When the fixed vector is one-hot, we only need to constrain the output to be one-hot, which however produces zero gradients almost everywhere and thus makes gradient-based optimization difficult. In this work, we introduce the sparse regularization strategy to approximate the one-hot constraint, which is composed of network output sharpening operation that enforces the output distribution of a network to be sharp and the $\ell_p$-norm ($p\le 1$) regularization that promotes the network output to be sparse. This simple approach guarantees the robustness of arbitrary loss functions while not hindering the fitting ability. Experimental results demonstrate that our method can significantly improve the performance of commonly-used loss functions in the presence of noisy labels and class imbalance, and outperform the state-of-the-art methods. The code is available at https://github.com/hitcszx/lnl_sr.


Applications of Artificial Neural Networks in Microorganism Image Analysis: A Comprehensive Review from Conventional Multilayer Perceptron to Popular Convolutional Neural Network and Potential Visual Transformer

arXiv.org Artificial Intelligence

Microorganisms are widely distributed in the human daily living environment. They play an essential role in environmental pollution control, disease prevention and treatment, and food and drug production. The identification, counting, and detection are the basic steps for making full use of different microorganisms. However, the conventional analysis methods are expensive, laborious, and time-consuming. To overcome these limitations, artificial neural networks are applied for microorganism image analysis. We conduct this review to understand the development process of microorganism image analysis based on artificial neural networks. In this review, the background and motivation are introduced first. Then, the development of artificial neural networks and representative networks are introduced. After that, the papers related to microorganism image analysis based on classical and deep neural networks are reviewed from the perspectives of different tasks. In the end, the methodology analysis and potential direction are discussed.