Deep Learning
Learning Accurate Integer Transformer Machine-Translation Models
We describe a method for training accurate Transformer machine-translation models to run inference using 8-bit integer (INT8) hardware matrix multipliers, as opposed to the more costly single-precision floating-point (FP32) hardware. Unlike previous work, which converted only 85 Transformer matrix multiplications to INT8, leaving 48 out of 133 of them in FP32 because of unacceptable accuracy loss, we convert them all to INT8 without compromising accuracy. Tested on the new-stest2014 English-to-German translation task, our INT8 Transformer Base and Transformer Big models yield BLEU scores that are 99.3% to 100% relative to those of the corresponding FP32 models. Our approach converts all matrix-multiplication tensors from an existing FP32 model into INT8 tensors by automatically making range-precision tradeoffs during training. To demonstrate the robustness of this approach, we also include results from INT6 Transformer models. 1 Introduction We report a method for training accurate yet compact Transformer machine-translation models [ V aswaniet al., 2017 ] . Specifically, we aim these models at hardware with 8-bit integer (INT8) matrix multipliers. Compared to single-precision floating-point (FP32) matrix multiplications, INT8 matrix multiplications not only reduce both storage and bandwidth four times, but they also consume 15 times less energy [ Horowitz, 2014 ] .
Optimizing Wireless Systems Using Unsupervised and Reinforced-Unsupervised Deep Learning
Liu, Dong, Sun, Chengjian, Yang, Chenyang, Hanzo, Lajos
Resource allocation and transceivers in wireless networks are usually designed by solving optimization problems subject to specific constraints, which can be formulated as variable or functional optimization. If the objective and constraint functions of a variable optimization problem can be derived, standard numerical algorithms can be applied for finding the optimal solution, which however incur high computational cost when the dimension of the variable is high. To reduce the on-line computational complexity, learning the optimal solution as a function of the environment's status by deep neural networks (DNNs) is an effective approach. DNNs can be trained under the supervision of optimal solutions, which however, is not applicable to the scenarios without models or for functional optimization where the optimal solutions are hard to obtain. If the objective and constraint functions are unavailable, reinforcement learning can be applied to find the solution of a functional optimization problem, which is however not tailored to optimization problems in wireless networks. In this article, we introduce unsupervised and reinforced-unsupervised learning frameworks for solving both variable and functional optimization problems without the supervision of the optimal solutions. When the mathematical model of the environment is completely known and the distribution of environment's status is known or unknown, we can invoke unsupervised learning algorithm. When the mathematical model of the environment is incomplete, we introduce reinforced-unsupervised learning algorithms that learn the model by interacting with the environment. Our simulation results confirm the applicability of these learning frameworks by taking a user association problem as an example.
On the comparability of Pre-trained Language Models
Aรenmacher, Matthias, Heumann, Christian
Recent developments in unsupervised representation learning have successfully established the concept of transfer learning in NLP. Mainly three forces are driving the improvements in this area of research: More elaborated architectures are making better use of contextual information. Instead of simply plugging in static pre-trained representations, these are learned based on surrounding context in end-to-end trainable models with more intelligently designed language modelling objectives. Along with this, larger corpora are used as resources for pre-training large language models in a self-supervised fashion which are afterwards fine-tuned on supervised tasks. Advances in parallel computing as well as in cloud computing, made it possible to train these models with growing capacities in the same or even in shorter time than previously established models. These three developments agglomerate in new state-of-the-art (SOTA) results being revealed in a higher and higher frequency. It is not always obvious where these improvements originate from, as it is not possible to completely disentangle the contributions of the three driving forces. We set ourselves to providing a clear and concise overview on several large pre-trained language models, which achieved SOTA results in the last two years, with respect to their use of new architectures and resources. We want to clarify for the reader where the differences between the models are and we furthermore attempt to gain some insight into the single contributions of lexical/computational improvements as well as of architectural changes. We explicitly do not intend to quantify these contributions, but rather see our work as an overview in order to identify potential starting points for benchmark comparisons. Furthermore, we tentatively want to point at potential possibilities for improvement in the field of open-sourcing and reproducible research.
Information Extraction based on Named Entity for Tourism Corpus
Chantrapornchai, Chantana, Tunsakul, Aphisit
Tourism information is scattered around nowadays. To search for the information, it is usually time consuming to browse through the results from search engine, select and view the details of each accommodation. In this paper, we present a methodology to extract particular information from full text returned from the search engine to facilitate the users. Then, the users can specifically look to the desired relevant information. The approach can be used for the same task in other domains. The main steps are 1) building training data and 2) building recognition model. First, the tourism data is gathered and the vocabularies are built. The raw corpus is used to train for creating vocabulary embedding. Also, it is used for creating annotated data. The process of creating named entity annotation is presented. Then, the recognition model of a given entity type can be built. From the experiments, given hotel description, the model can extract the desired entity,i.e, name, location, facility. The extracted data can further be stored as a structured information, e.g., in the ontology format, for future querying and inference. The model for automatic named entity identification, based on machine learning, yields the error ranging 8%-25%.
In search for Alzheimer's disease in the retina with AI - AIMed
"Eyes are the windows to the soul". It's probably many physicians' dreams to be able to tell what a patient has come down with by looking into their eyes. Researchers from the University College London (UCL) and Moorfields Eye Hospital are trying to realize this dream in a collaborative project called "AlzEye". By studying a database of eye scans which include details of patients' retina alongside with other vital health information, the research team hope to detect optical differences and see if they may be telltale signs of Alzheimer's disease. To facilitate the process, the team is engaging with Google DeepMind, to employ machine learning algorithms to go through scans and information of 300,000 patients aged 40 and above who had visited Moorfields between year 2008 and 2018.
Are We Overly Infatuated With Deep Learning? - CTOvision.com
One of the factors often credited for this latest boom in artificial intelligence (AI) investment, research, and related cognitive technologies, is the emergence of deep learning neural networks as an evolution of machine algorithms, as well as the corresponding large volume of big data and computing power that makes deep learning a practical reality. While deep learning has been extremely popular and has shown real ability to solve many machine learning problems, deep learning is just one approach to machine learning (ML), that while having proven much capability across a wide range of problem areas, is still just one of many practical approaches.
Can MRI predict intelligence levels in children?
A group of researchers from the Skoltech Center for Computational and Data-Intensive Science and Engineering (CDISE) took 4th place in the international MRI-based adolescent intelligence prediction competition. For the first time ever, the Skoltech scientists used ensemble methods based on deep learning 3-D networks to deal with this challenging prediction task. The results of their study were published in the journal Adolescent Brain Cognitive Development Neurocognitive Prediction. In 2013, the US National Institutes of Health (NIH) launched the first grand-scale study of its kind in adolescent brain research, Adolescent Brain Cognitive Development (ABCD, abcdstudy.org/), Magnetic Resonance Imaging (MRI) is a common technique used to obtain images of human internal organs and tissues.
New DeepMind AI 'spots breast cancer better than clinicians'
A newly developed artificial intelligence (AI) model is able to spot breast cancer better than a clinician, new research has suggested. Google DeepMind, in partnership with Cancer Research UK Imperial Centre, Northwestern University and Royal Surrey County Hospital, has developed the model which can spot cancer in breast screening mammograms in a bid to improve health outcomes and ease pressure on overstretched radiology services. Initial findings, published by the technology giant in the journal Nature, suggest the AI can identify the disease with greater accuracy, fewer false positives and fewer false negatives. The model, trained on de-identified data of 76,000 women in the UK and more than 15,000 women in the US, reportedly lowered false positive results by 1.2% and false negatives by 2.7% in the UK, but is yet to be tested in clinical studies. When tested, the AI system processed only the latest available mammogram of a patient, whereas clinicians had access to patient histories and prior mammograms to make an informed screening decision.
Happy AI New Year! Global Researchers Reflect on 2019, Talk Trends for 2020
The year 2019 saw unprecedented growth in AI research, development and deployment. Great technical progress has been achieved in image recognition, image generation, natural language understanding and other fields; while challenges remain with data management, efficiency measurement, computational capacity and other issues. To welcome 2020 with some fresh AI perspectives, Synced spoke with global researchers from Google Brain, Sony AI, Alibaba affiliate Ant Financial (formerly known as Alipay), Israel-based AI processor company Habana (recently acquired by Intel), Russian tech giant Yandex, Vietnam's newly established research lab VinAI Research, French deep learning inference acceleration startup Mipsology, and China-based remote sensing data platform TerraQuanta. Colin Raffel, Senior Research Scientist, Google Brain In 2019 the community made huge progress on learning from limited labels. MixMatch, UDA, S4L, and ReMixMatch produced huge gains on standard semi-supervised learning benchmarks.
Artificial Neural Networks For Blockchain: A Primer
Artificial neural networks (ANNs) have proven to be extremely useful for solving problems such as classification, regression, function estimation and dimensionality reduction. However, it turns out that different neural network architectures are able to achieve higher performances for certain problems. This article will provide an overview of the most common neural network architectures -- including recurrent neural networks and convolutional neural -- and how they can be implemented to aid blockchain technology. Convolutional neural networks (CNNs) are a type of neural network that is designed to capture increasingly more complex features within its input data. To do this, CNNs are constructed from a sequence of layers, each of which consists of a series of cube-shaped filters.