Deep Learning
Left Ventricle Contouring in Cardiac Images Based on Deep Reinforcement Learning
Yin, Sixing, Han, Yameng, Li, Shufang
Medical image segmentation is one of the important tasks of computer-aided diagnosis in medical image analysis. Since most medical images have the characteristics of blurred boundaries and uneven intensity distribution, through existing segmentation methods, the discontinuity within the target area and the discontinuity of the target boundary are likely to lead to rough or even erroneous boundary delineation. In this paper, we propose a new iterative refined interactive segmentation method for medical images based on agent reinforcement learning, which focuses on the problem of target segmentation boundaries. We model the dynamic process of drawing the target contour in a certain order as a Markov Decision Process (MDP) based on a deep reinforcement learning method. In the dynamic process of continuous interaction between the agent and the image, the agent tracks the boundary point by point in order within a limited length range until the contour of the target is completely drawn. In this process, the agent can quickly improve the segmentation performance by exploring an interactive policy in the image. The method we proposed is simple and effective. At the same time, we evaluate our method on the cardiac MRI scan data set. Experimental results show that our method has a better segmentation effect on the left ventricle in a small number of medical image data sets, especially in terms of segmentation boundaries, this method is better than existing methods. Based on our proposed method, the dynamic generation process of the predicted contour trajectory of the left ventricle will be displayed online at https://github.com/H1997ym/LV-contour-trajectory.
Position Bias Mitigation: A Knowledge-Aware Graph Model for Emotion Cause Extraction
Yan, Hanqi, Gui, Lin, Pergola, Gabriele, He, Yulan
The Emotion Cause Extraction (ECE)} task aims to identify clauses which contain emotion-evoking information for a particular emotion expressed in text. We observe that a widely-used ECE dataset exhibits a bias that the majority of annotated cause clauses are either directly before their associated emotion clauses or are the emotion clauses themselves. Existing models for ECE tend to explore such relative position information and suffer from the dataset bias. To investigate the degree of reliance of existing ECE models on clause relative positions, we propose a novel strategy to generate adversarial examples in which the relative position information is no longer the indicative feature of cause clauses. We test the performance of existing models on such adversarial examples and observe a significant performance drop. To address the dataset bias, we propose a novel graph-based method to explicitly model the emotion triggering paths by leveraging the commonsense knowledge to enhance the semantic dependencies between a candidate clause and an emotion clause. Experimental results show that our proposed approach performs on par with the existing state-of-the-art methods on the original ECE dataset, and is more robust against adversarial attacks compared to existing models.
Transient Chaos in BERT
Inoue, Katsuma, Ohara, Soh, Kuniyoshi, Yasuo, Nakajima, Kohei
Language is an outcome of our complex and dynamic human-interactions and the technique of natural language processing (NLP) is hence built on human linguistic activities. Bidirectional Encoder Representations from Transformers (BERT) has recently gained its popularity by establishing the state-of-the-art scores in several NLP benchmarks. A Lite BERT (ALBERT) is literally characterized as a lightweight version of BERT, in which the number of BERT parameters is reduced by repeatedly applying the same neural network called Transformer's encoder layer. By pre-training the parameters with a massive amount of natural language data, ALBERT can convert input sentences into versatile high-dimensional vectors potentially capable of solving multiple NLP tasks. In that sense, ALBERT can be regarded as a well-designed high-dimensional dynamical system whose operator is the Transformer's encoder, and essential structures of human language are thus expected to be encapsulated in its dynamics. In this study, we investigated the embedded properties of ALBERT to reveal how NLP tasks are effectively solved by exploiting its dynamics. We thereby aimed to explore the nature of human language from the dynamical expressions of the NLP model. Our short-term analysis clarified that the pre-trained model stably yields trajectories with higher dimensionality, which would enhance the expressive capacity required for NLP tasks. Also, our long-term analysis revealed that ALBERT intrinsically shows transient chaos, a typical nonlinear phenomenon showing chaotic dynamics only in its transient, and the pre-trained ALBERT model tends to produce the chaotic trajectory for a significantly longer time period compared to a randomly-initialized one. Our results imply that local chaoticity would contribute to improving NLP performance, uncovering a novel aspect in the role of chaotic dynamics in human language behaviors.
Spatial Graph Attention and Curiosity-driven Policy for Antiviral Drug Discovery
Wu, Yulun, Choma, Nicholas, Chen, Andrew, Cashman, Mikaela, Prates, รrica T., Shah, Manesh, Vergara, Verรณnica G. Melesse, Clyde, Austin, Brettin, Thomas S., de Jong, Wibe A., Kumar, Neeraj, Head, Martha S., Stevens, Rick L., Nugent, Peter, Jacobson, Daniel A., Brown, James B.
We developed Distilled Graph Attention Policy Networks (DGAPNs), a curiosity-driven reinforcement learning model to generate novel graph-structured chemical representations that optimize user-defined objectives by efficiently navigating a physically constrained domain. The framework is examined on the task of generating molecules that are designed to bind, noncovalently, to functional sites of SARS-CoV-2 proteins. We present a spatial Graph Attention Network (sGAT) that leverages self-attention over both node and edge attributes as well as encoding spatial structure -- this capability is of considerable interest in areas such as molecular and synthetic biology and drug discovery. An attentional policy network is then introduced to learn decision rules for a dynamic, fragment-based chemical environment, and state-of-the-art policy gradient techniques are employed to train the network with enhanced stability. Exploration is efficiently encouraged by incorporating innovation reward bonuses learned and proposed by random network distillation. In experiments, our framework achieved outstanding results compared to state-of-the-art algorithms, while increasing the diversity of proposed molecules and reducing the complexity of paths to chemical synthesis.
Objective Robustness in Deep Reinforcement Learning
Koch, Jack, Langosco, Lauro, Pfau, Jacob, Le, James, Sharkey, Lee
We study objective robustness failures, a type of out-of-distribution robustness failure in reinforcement learning (RL). Objective robustness failures occur when an RL agent retains its capabilities out-of-distribution yet pursues the wrong objective. This kind of failure presents different risks than the robustness problems usually considered in the literature, since it involves agents that leverage their capabilities to pursue the wrong objective rather than simply failing to do anything useful. We provide the first explicit empirical demonstrations of objective robustness failures and present a partial characterization of its causes.
Harmless Overparametrization in Two-layer Neural Networks
In such a wide range of applications, neural networks are preferred to be overparametrized in the sense that the number of active (nonzero) network parameters is much larger than the sample size. One example is the Alex-Net (Krizhevsky, Sutskever and Hinton, 2012). There is convincing evidence showing that overparametrization can help optimization (Arora, Cohen and Hazan, 2018; Safran, Yehudai and Shamir, 2020), however, overparametrized deep neural networks can easily fit random labels even in the presence of explicit regularization (Zhang et al., 2016). This indicates that overparametrized models have the capability to overfit, but do not necessarily lead to bad testing performance. Thus two intriguing questions come out: why can overparametrized neural networks exhibit good testing performance and how good can the testing performance be? Overparametrization, where the number of active parameters is larger than the sample size, is not new in statistics. For example, high dimensional linear models can be viewed to be overparametrized. In high dimensional linear regression, prediction is closely related to the estimation of the true parameter and overparametrization can be harmful to estimation even in the presence of optimal regularization. For example, the prediction risk of the optimal regularized ridge estimator can have a strictly positive limit in the overparametrized settings (Dobriban et al., 2018; Hastie et al., 2019).
Self-Supervised Learning with Data Augmentations Provably Isolates Content from Style
von Kรผgelgen, Julius, Sharma, Yash, Gresele, Luigi, Brendel, Wieland, Schรถlkopf, Bernhard, Besserve, Michel, Locatello, Francesco
Self-supervised representation learning has shown remarkable success in a number of domains. A common practice is to perform data augmentation via hand-crafted transformations intended to leave the semantics of the data invariant. We seek to understand the empirical success of this approach from a theoretical perspective. We formulate the augmentation process as a latent variable model by postulating a partition of the latent representation into a content component, which is assumed invariant to augmentation, and a style component, which is allowed to change. Unlike prior work on disentanglement and independent component analysis, we allow for both nontrivial statistical and causal dependencies in the latent space. We study the identifiability of the latent representation based on pairs of views of the observations and prove sufficient conditions that allow us to identify the invariant content partition up to an invertible mapping in both generative and discriminative settings. We find numerical simulations with dependent latent variables are consistent with our theory. Lastly, we introduce Causal3DIdent, a dataset of high-dimensional, visually complex images with rich causal dependencies, which we use to study the effect of data augmentations performed in practice.
A self consistent theory of Gaussian Processes captures feature learning effects in finite CNNs
Deep neural networks (DNNs) in the infinite width/channel limit have received much attention recently, as they provide a clear analytical window to deep learning via mappings to Gaussian Processes (GPs). Despite its theoretical appeal, this viewpoint lacks a crucial ingredient of deep learning in finite DNNs, laying at the heart of their success -- feature learning. Here we consider DNNs trained with noisy gradient descent on a large training set and derive a self consistent Gaussian Process theory accounting for strong finite-DNN and feature learning effects. Applying this to a toy model of a two-layer linear convolutional neural network (CNN) shows good agreement with experiments. We further identify, both analytical and numerically, a sharp transition between a feature learning regime and a lazy learning regime in this model. Strong finite-DNN effects are also derived for a non-linear two-layer fully connected network. Our self consistent theory provides a rich and versatile analytical framework for studying feature learning and other non-lazy effects in finite DNNs.
Natural Language Processing for Stocks News Analysis
In this hands-on project, we will train a Long Short Term Memory (LSTM) deep learning model to perform stocks sentiment analysis. Natural language processing (NLP) works by converting words (text) into numbers, these numbers are then used to train an AI/ML model to make predictions. In this project, we will build a machine learning model to analyze thousands of Twitter tweets to predict people's sentiment towards a particular company or stock. The algorithm could be used automatically understand the sentiment from public tweets, which could be used as a factor while making buy/sell decision of securities. Note: This course works best for learners who are based in the North America region.
Why Machines Don't Speak Spanish Well (and Why They Should)
Every day people talk more naturally about artificial intelligence (AI). We are getting used to this label - with a meaning for many still surrounded by an enigmatic halo - penetrating our routine more frequently. Without being barely conscious, we smile to unlock the mobile phone without knowing that after that second in front of the camera, thousands of pixels converted into data feed deep learning algorithms at high speed. These are today capable of automating facial recognition in percentages greater than 98% accuracy. The hatching has been stellar.