Machine Translation
Natural Language Generation for Electronic Health Records
A variety of methods existing for generating synthetic electronic health records (EHRs), but they are not capable of generating unstructured text, like emergency department (ED) chief complaints, history of present illness or progress notes. Here, we use the encoder-decoder model, a deep learning algorithm that features in many contemporary machine translation systems, to generate synthetic chief complaints from discrete variables in EHRs, like age group, gender, and discharge diagnosis. After being trained end-to-end on authentic records, the model can generate realistic chief complaint text that preserves much of the epidemiological information in the original data. As a side effect of the model's optimization goal, these synthetic chief complaints are also free of relatively uncommon abbreviation and misspellings, and they include none of the personally-identifiable information (PII) that was in the training data, suggesting it may be used to support the de-identification of text in EHRs. When combined with algorithms like generative adversarial networks (GANs), our model could be used to generate fully-synthetic EHRs, facilitating data sharing between healthcare providers and researchers and improving our ability to develop machine learning methods tailored to the information in healthcare data. 1 Introduction The wide adoption of electronic health record (EHR) systems has led to the creation of large amounts of healthcare data. Although these data are primarily used to improve patient outcomes and streamline the delivery of care (healthit.gov), Because they contain personally identifiable patient information, however, much of which is protected under the Health Insurance Portability and Accountability Act (HIPAA), these data are often difficult for providers to share with investigators outside their organizations, limiting their feasibility for use in research.
A Stochastic Decoder for Neural Machine Translation
Schulz, Philip, Aziz, Wilker, Cohn, Trevor
The process of translation is ambiguous, in that there are typically many valid trans- lations for a given sentence. This gives rise to significant variation in parallel cor- pora, however, most current models of machine translation do not account for this variation, instead treating the prob- lem as a deterministic process. To this end, we present a deep generative model of machine translation which incorporates a chain of latent variables, in order to ac- count for local lexical and syntactic varia- tion in parallel corpora. We provide an in- depth analysis of the pitfalls encountered in variational inference for training deep generative models. Experiments on sev- eral different language pairs demonstrate that the model consistently improves over strong baselines.
Reliability and Learnability of Human Bandit Feedback for Sequence-to-Sequence Reinforcement Learning
Kreutzer, Julia, Uyheng, Joshua, Riezler, Stefan
We present a study on reinforcement learning (RL) from human bandit feedback for sequence-to-sequence learning, exemplified by the task of bandit neural machine translation (NMT). We investigate the reliability of human bandit feedback, and analyze the influence of reliability on the learnability of a reward estimator, and the effect of the quality of reward estimates on the overall RL task. Our analysis of cardinal (5-point ratings) and ordinal (pairwise preferences) feedback shows that their intra- and inter-annotator $\alpha$-agreement is comparable. Best reliability is obtained for standardized cardinal feedback, and cardinal feedback is also easiest to learn and generalize from. Finally, improvements of over 1 BLEU can be obtained by integrating a regression-based reward estimator trained on cardinal feedback for 800 translations into RL for NMT. This shows that RL is possible even from small amounts of fairly reliable human feedback, pointing to a great potential for applications at larger scale.
Deep Graph Translation
Guo, Xiaojie, Wu, Lingfei, Zhao, Liang
Inspired by the tremendous success of deep generative models on generating continuous data like image and audio, in the most recent year, few deep graph generative models have been proposed to generate discrete data such as graphs. They are typically unconditioned generative models which has no control on modes of the graphs being generated. Differently, in this paper, we are interested in a new problem named \emph{Deep Graph Translation}: given an input graph, we want to infer a target graph based on their underlying (both global and local) translation mapping. Graph translation could be highly desirable in many applications such as disaster management and rare event forecasting, where the rare and abnormal graph patterns (e.g., traffic congestions and terrorism events) will be inferred prior to their occurrence even without historical data on the abnormal patterns for this graph (e.g., a road network or human contact network). To achieve this, we propose a novel Graph-Translation-Generative Adversarial Networks (GT-GAN) which will generate a graph translator from input to target graphs. GT-GAN consists of a graph translator where we propose new graph convolution and deconvolution layers to learn the global and local translation mapping. A new conditional graph discriminator has also been proposed to classify target graphs by conditioning on input graphs. Extensive experiments on multiple synthetic and real-world datasets demonstrate the effectiveness and scalability of the proposed GT-GAN.
Refining Source Representations with Relation Networks for Neural Machine Translation
Zhang, Wen, Hu, Jiawei, Feng, Yang, Liu, Qun
Although neural machine translation (NMT) with the encoder-decoder framework has achieved great success in recent times, it still suffers from some drawbacks: RNNs tend to forget old information which is often useful in the current step and the encoder only operates over words without considering word relationship. To solve these problems, we introduce relation networks (RNs) to learn better representations of the source. In our method RNs are used to associate source words with each other so that the source representation can memorize all the source words and also contain the relationship between them. Then the source representations and all the relations are fed into the attention component together while decoding, with the main encoder-decoder architecture unchanged. Experiments on several data sets show that our method can improve the translation performance significantly over the conventional encoder-decoder model, and can even outperform the approach involving supervised syntactic knowledge.
Refining Source Representations with Relation Networks for Neural Machine Translation
Zhang, Wen, Hu, Jiawei, Feng, Yang, Liu, Qun
Although neural machine translation (NMT) with the encoder-decoder framework has achieved great success in recent times, it still suffers from some drawbacks: RNNs tend to forget old information which is often useful and the encoder only operates through words without considering word relationship. To solve these problems, we introduce a relation networks (RN) into NMT to refine the encoding representations of the source. In our method, the RN first augments the representation of each source word with its neighbors and reasons all the possible pairwise relations between them. Then the source representations and all the relations are fed to the attention module and the decoder together, keeping the main encoder-decoder architecture unchanged. Experiments on two Chinese-to-English data sets in different scales both show that our method can outperform the competitive baselines significantly.
NVIDIA's AI-Driven Data Center Business Could Grow by 18 Times in 5 Years
NVIDIA (NASDAQ:NVDA) recently reported powerful fiscal first-quarter 2018 results. The graphics processing unit (GPU) specialist's revenue jumped 66%, GAAP earnings per share soared 151%, and adjusted EPS surged 141%. There was a wealth of information about the company's results and future prospects shared on the earnings call. Our focus here is on NVIDIA's data center, which is growing like gangbusters -- its revenue grew 71% year over year to $701 million in the quarter, accounting for 22% of the company's total revenue. We see the data center opportunity as very large, fueled by growing demand for accelerated computing and applications ranging from AI [artificial intelligence] to high-performance computing across multiple market segments and vertical industries.
How AI Is Making Prediction Cheaper
Avi Goldfarb, a professor at the University of Toronto's Rotman School of Management, explains the economics of machine learning, a branch of artificial intelligence that makes predictions. He says as prediction gets cheaper and better, machines are going to be doing more of it. That means businesses -- and individual workers -- need to figure out how to take advantage of the technology to stay competitive. Goldfarb is the coauthor of the book Prediction Machines: The Simple Economics of Artificial Intelligence. CURT NICKISCH: Welcome to the HBR IdeaCast, from Harvard Business Review. YOUTUBE: [Two women speaking] We've got this all tabbed up? In it, three young English-speaking women use Google Translate to order food in Hindi from an Indian restaurant. They copy and paste their order in English into the computer, and it translates items like "samosas" and reads them aloud in the foreign language.
NVIDIA's AI-Driven Data Center Business Could Grow by 18 Times in 5 Years
NVIDIA (NASDAQ: NVDA) recently reported powerful fiscal first-quarter 2018 results. The graphics processing unit (GPU) specialist's revenue jumped 66%, GAAP earnings per share soared 151%, and adjusted EPS surged 141%. There was a wealth of information about the company's results and future prospects shared on the earnings call. Our focus here is on NVIDIA's data center, which is growing like gangbusters -- its revenue grew 71% year over year to $701 million in the quarter, accounting for 22% of the company's total revenue. We see the data center opportunity as very large, fueled by growing demand for accelerated computing and applications ranging from AI [artificial intelligence] to high-performance computing across multiple market segments and vertical industries.
A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings
Artetxe, Mikel, Labaka, Gorka, Agirre, Eneko
Recent work has managed to learn cross-lingual word embeddings without parallel data by mapping monolingual embeddings to a shared space through adversarial training. However, their evaluation has focused on favorable conditions, using comparable corpora or closely-related languages, and we show that they often fail in more realistic scenarios. This work proposes an alternative approach based on a fully unsupervised initialization that explicitly exploits the structural similarity of the embeddings, and a robust self-learning algorithm that iteratively improves this solution. Our method succeeds in all tested scenarios and obtains the best published results in standard datasets, even surpassing previous supervised systems. Our implementation is released as an open source project at https://github.com/artetxem/vecmap