Deep Learning
Continuous Learning of Context-dependent Processing in Neural Networks
Zeng, Guanxiong, Chen, Yang, Cui, Bo, Yu, Shan
Deep artificial neural networks (DNNs) are powerful tools for recognition and classification as they learn sophisticated mapping rules between the inputs and the outputs. However, the rules that learned by the majority of current DNNs used for pattern recognition are largely fixed and do not vary with different conditions. This limits the network's ability to work in more complex and dynamical situations in which the mapping rules themselves are not fixed but constantly change according to contexts, such as different environments and goals. Inspired by the role of the prefrontal cortex (PFC) in mediating context-dependent processing in the primate brain, here we propose a novel approach, involving a learning algorithm named orthogonal weights modification (OWM) with the addition of a PFC-like module, that enables networks to continually learn different mapping rules in a context-dependent way. We demonstrate that with OWM to protect previously acquired knowledge, the networks could sequentially learn up to thousands of different mapping rules without interference, and needing as few as $\sim$10 samples to learn each, reaching a human level ability in online, continual learning. In addition, by using a PFC-like module to enable contextual information to modulate the representation of sensory features, a network could sequentially learn different, context-specific mappings for identical stimuli. Taken together, these approaches allow us to teach a single network numerous context-dependent mapping rules in an online, continual manner. This would enable highly compact systems to gradually learn myriad of regularities of the real world and eventually behave appropriately within it.
AdaShift: Decorrelation and Convergence of Adaptive Learning Rate Methods
Zhou, Zhiming, Zhang, Qingru, Lu, Guansong, Wang, Hongwei, Zhang, Weinan, Yu, Yong
Adam is shown not being able to converge to the optimal solution in certain cases. Researchers recently propose several algorithms to avoid the issue of non-convergence of Adam, but their efficiency turns out to be unsatisfactory in practice. In this paper, we provide a new insight into the non-convergence issue of Adam as well as other adaptive learning rate methods. We argue that there exists an inappropriate correlation between gradient $g_t$ and the second moment term $v_t$ in Adam ($t$ is the timestep), which results in that a large gradient is likely to have small step size while a small gradient may have a large step size. We demonstrate that such unbalanced step sizes are the fundamental cause of non-convergence of Adam, and we further prove that decorrelating $v_t$ and $g_t$ will lead to unbiased step size for each gradient, thus solving the non-convergence problem of Adam. Finally, we propose AdaShift, a novel adaptive learning rate method that decorrelates $v_t$ and $g_t$ by temporal shifting, i.e., using temporally shifted gradient $g_{t-n}$ to calculate $v_t$. The experiment results demonstrate that AdaShift is able to address the non-convergence issue of Adam, while still maintaining a competitive performance with Adam in terms of both training speed and generalization.
Computationally Efficient Cascaded Training for Deep Unrolled Network in CT Imaging
Wu, Dufan, Kim, Kyungsang, Li, Quanzheng
Abstract--Dose reduction in computed tomography (CT) has been of great research interest for decades with the endeavor to reduce the health risk related to radiation. Promising results have been achieved by the recent application of deep learning to image reconstruction algorithms. Unrolled neural networks have reached state-of-the-art performance by learning the image reconstruction algorithm end-to-end. However, it suffers from huge memory consumption and long training time, which made it hard to scale to 3D data with current hardware. In this paper, we proposed an unrolled neural network for image reconstruction which can be trained step-by-step instead of end-to-end. Multiple cascades of image domain network were trained sequentially and connected with iterations which enforced data fidelity. Local image patches could be utilized for the neural network training, which made it fully scalable to 3D CT data. The proposed method was validated with both simulated and real data and demonstrated competing performance against the end-to-end networks. Tube current reduction is currently the most practical way to achieve lower-dose scans [3]. Meanwhile, sparse-view sampling has also demonstrated great potential in dose reduction, and corresponding systems are under active development [4].
A Deep Autoencoder System for Differentiation of Cancer Types Based on DNA Methylation State
Khwaja, Mohammed, Kalofonou, Melpomeni, Toumazou, Chris
Abstract--A Deep Autoencoder based content retrieval algorithm is proposed for prediction and differentiation of cancer types based on the presence of epigenetic patterns of DNA methylation identified in genetic regions known as CpG islands. The developed deep learning system uses a CpG island state classification subsystem to complete sets of missing/incomplete island data in given human cell lines, and is then pipelined with an intricate set of statistical and signal processing methods to accurately predict the presence of cancer and further differentiate the type and cell of origin in the event of a positive result. The proposed system was trained with previously reported data derived from four case groups of cancer cell lines, achieving overall Sensitivity of 88.24%, Specificity of 83.33%, Accuracy of 84.75% and Matthews Correlation Coefficient of 0.687. The ability to predict and differentiate cancer types using epigenetic events as the identifying patterns was demonstrated in previously reported data sets from breast, lung, lymphoblastic leukemia and urological cancer cell lines, allowing the pipelined system to be robust and adjustable to other cancer cell lines or epigenetic events. Significant progress has been made in understanding crucial regulatory mechanisms responsible for the development and progression of cancer at a cellular and molecular level, through genetic alterations such as DNA mutations and disruptions in epigenetic mechanisms including DNA methylation and histone modifications [1]. Cancer rates have been progressively increasing, with the latest statistics from Cancer Research UK to have reported more than 350,000 new cases diagnosed in the UK [2], of which more than 40% could have been prevented. Cancer research has been significantly progressing with advances in more effective treatments and screening methods, however there is still a pressing need for more targeted methods to be available for monitoring of cancer progression and prevention of treatment resistance that would help control the disease and improve survival rates.
Learning Neuron Non-Linearities with Kernel-Based Deep Neural Networks
Marra, Giuseppe, Zanca, Dario, Betti, Alessandro, Gori, Marco
The effectiveness of deep neural architectures has been widely supported in terms of both experimental and foundational principles. There is also clear evidence that the activation function (e.g. the rectifier and the LSTM units) plays a crucial role in the complexity of learning. Based on this remark, this paper discusses an optimal selection of the neuron non-linearity in a functional framework that is inspired from classic regularization arguments. It is shown that the best activation function is represented by a kernel expansion in the training set, that can be effectively approximated over an opportune set of points modeling 1-D clusters. The idea can be naturally extended to recurrent networks, where the expressiveness of kernel-based activation functions turns out to be a crucial ingredient to capture long-term dependencies. We give experimental evidence of this property by a set of challenging experiments, where we compare the results with neural architectures based on state of the art LSTM cells.
Neural Generation of Diverse Questions using Answer Focus, Contextual and Linguistic Features
Harrison, Vrindavan, Walker, Marilyn
Question Generation is the task of automatically creating questions from textual input. In this work we present a new Attentional Encoder--Decoder Recurrent Neural Network model for automatic question generation. Our model incorporates linguistic features and an additional sentence embedding to capture meaning at both sentence and word levels. The linguistic features are designed to capture information related to named entity recognition, word case, and entity coreference resolution. In addition our model uses a copying mechanism and a special answer signal that enables generation of numerous diverse questions on a given sentence. Our model achieves state of the art results of 19.98 Bleu_4 on a benchmark Question Generation dataset, outperforming all previously published results by a significant margin. A human evaluation also shows that these added features improve the quality of the generated questions.
Hybrid Active Inference
Ofner, Andrรฉ, Stober, Sebastian
We describe a framework of hybrid cognition by formulating a hybrid cognitive agent that performs hierarchical active inference across a human and a machine part. We suggest that, in addition to enhancing human cognitive functions with an intelligent and adaptive interface, integrated cognitive processing could accelerate emergent properties within artificial intelligence. To establish this, a machine learning part learns to integrate into human cognition by explaining away multi-modal sensory measurements from the environment and physiology simultaneously with the brain signal. With ongoing training, the amount of predictable brain signal increases. This lends the agent the ability to self-supervise on increasingly high levels of cognitive processing in order to further minimize surprise in predicting the brain signal. Furthermore, with increasing level of integration, the access to sensory information about environment and physiology is substituted with access to their representation in the brain. While integrating into a joint embodiment of human and machine, human action and perception are treated as the machine's own. The framework can be implemented with invasive as well as non-invasive sensors for environment, body and brain interfacing. Online and offline training with different machine learning approaches are thinkable. Building on previous research on shared representation learning, we suggest a first implementation leading towards hybrid active inference with non-invasive brain interfacing and state of the art probabilistic deep learning methods. We further discuss how implementation might have effect on the meta-cognitive abilities of the described agent and suggest that with adequate implementation the machine part can continue to execute and build upon the learned cognitive processes autonomously.
Model-Ensemble Trust-Region Policy Optimization
Kurutach, Thanard, Clavera, Ignasi, Duan, Yan, Tamar, Aviv, Abbeel, Pieter
Model-free reinforcement learning (RL) methods are succeeding in a growing number of tasks, aided by recent advances in deep learning. However, they tend to suffer from high sample complexity, which hinders their use in real-world domains. Alternatively, model-based reinforcement learning promises to reduce sample complexity, but tends to require careful tuning and to date have succeeded mainly in restrictive domains where simple models are sufficient for learning. In this paper, we analyze the behavior of vanilla model-based reinforcement learning methods when deep neural networks are used to learn both the model and the policy, and show that the learned policy tends to exploit regions where insufficient data is available for the model to be learned, causing instability in training. To overcome this issue, we propose to use an ensemble of models to maintain the model uncertainty and regularize the learning process. We further show that the use of likelihood ratio derivatives yields much more stable learning than backpropagation through time. Altogether, our approach Model-Ensemble Trust-Region Policy Optimization (ME-TRPO) significantly reduces the sample complexity compared to model-free deep RL methods on challenging continuous control benchmark tasks.
How AI makes recruiting more human (via Passle)
Artificial intelligence, machine learning and deep learning have dominated industries and headlines for years. This begs the question: how will AI affect me and my role? The recruiting industry has been particularly susceptible, as AI and advanced technology can parse through an avalanche of resumes faster than you can Snapchat your breakfast. It thrives in key processes such as data generation, development of algorithms and the evaluation of results. AI can look into the vast universe of social media and beyond, detect great prospective employees and understand their skills and capabilities on a deeper level, which may be easily missed on a standard resume.
Detecting Fake News, At Its Source
Summary: Researchers have created a new deep learning system that can determine if a news outlet is accurate or biased based on only 150 articles published. The algorithm can also detect the political leanings of a news site. Researchers say fake news articles are more likely to use language that is hyperbolic, subjective and emotional.