Machine Translation
Where can I find a trained machine translation seq2seq model? • /r/MachineLearning
Where can I find a trained machine translation seq2seq model? Title says it all - I'd like to play around with a well trained LSTM sequence to sequence MT model, but I'd rather not futz around with training one. I am not aware of any that exist. You will have to train one yourself or convince someone else to train one for you. If you just want to push input and get output, you might find this demo from Bengio's lab interesting: http://104.131.78.120/
Machine Translation: The Combination of Machine Learning and Human Intelligence - insideBIGDATA
In this special guest feature, Vasco Pedro, CEO and Co-Founder of Unbabel, discusses the importance of machine translation for natural languages and how it currently lacks the quality companies demand for their content. Dr. Pedro' company is Unbabel, the Y Combinator-backed startup that combines crowdsourced human translation and machine learning to deliver fast translation services to businesses with human tone and nuance. Vasco previously worked for Google helping to develop technology for data computation and language at scale, and served as a research faculty member at the Technical University of Lisbon. Vasco holds a PhD in Language Technologies from Carnegie Mellon University in the field of computational semantics. Additionally, Vasco is a Fulbright Scholar, mentor, and advisor to a number of startups on top of being a serial entrepreneur.
Hands-free speech translation app gets trialed at Narita airport
The "NariTra" multilingual translation app employs noise-canceling techniques and recognizes a wide range of speech. Offered by the airport at no cost, the app is designed to work hands-free -- and therefore suitable for foreign visitors who have just arrived and who have their hands full with luggage. The tests will see the app deployed on a shuttle bus running between Terminal 1 and Terminal 2, translating Japanese into English, Chinese and Korean, and vice versa. The airport operator plans to roll out the app by the time the 2020 Tokyo Olympic and Paralympic Games take place.
Text Simplification Using Neural Machine Translation
Wang, Tong (University of Massachusetts Boston) | Chen, Ping (University of Massachusetts Boston) | Rochford, John (University of Masschusetts Medical School) | Qiang, Jipeng (Hefei University of Technology)
Text simplification (TS) is the technique of reducing the lexical, syntactical complexity of text. Existing automatic TS systems can simplify text only by lexical simplification or by manually defined rules. Neural Machine Translation (NMT) is a recently proposed approach for Machine Translation (MT) that is receiving a lot of research interest. In this paper, we regard original English and simplified English as two languages, and apply a NMT model–Recurrent Neural Network (RNN) encoder-decoder on TS to make the neural network to learn text simplification rules by itself. Then we discuss challenges and strategies about how to apply a NMT model to the task of text simplification.
Estimating Text Intelligibility via Information Packaging Analysis
Li, Junyi Jessy (University of Pennsylvania)
Effective communication through language involves organizing the content a person or system wishes to convey into text that flows naturally. There are many ways to render the same information, but those appropriate for one group of audience may not be intelligible to another. The goal of this thesis to analyze and address factors that influence the intelligibility of text from two aspects of information packaging: discourse structure and text specificity. Effective communication through language involves organizing the content a person or system wishes to convey into text that flows naturally. There are many ways to render the same information, but those appropriate for one group of audience may not be intelligible to another. The goal of this thesis to analyze and address factors that influence the intelligibility of text from two aspects of information packaging: discourse structure and text specificity.
To Swap or Not to Swap? Exploiting Dependency Word Pairs for Reordering in Statistical Machine Translation
Hadiwinoto, Christian (National University of Singapore) | Liu, Yang (Tsinghua University) | Ng, Hwee Tou (National University of Singapore)
Reordering poses a major challenge in machine translation (MT) between two languages with significant differences in word order. In this paper, we present a novel reordering approach utilizing sparse features based on dependency word pairs. Each instance of these features captures whether two words, which are related by a dependency link in the source sentence dependency parse tree, follow the same order or are swapped in the translation output. Experiments on Chinese-to-English translation show a statistically significant improvement of 1.21 BLEU point using our approach, compared to a state-of-the-art statistical MT system that incorporates prior reordering approaches.
Building Earth Mover's Distance on Bilingual Word Embeddings for Machine Translation
Zhang, Meng (Tsinghua University) | Liu, Yang (Tsinghua University) | Luan, Huanbo (Tsinghua University) | Sun, Maosong (Tsinghua University) | Izuha, Tatsuya (Toshiba Corporation Corporate Research &) | Hao, Jie (Development Center)
Following their monolingual counterparts, bilingual word embeddings are also on the rise. As a major application task, word translation has been relying on the nearest neighbor to connect embeddings cross-lingually. However, the nearest neighbor strategy suffers from its inherently local nature and fails to cope with variations in realistic bilingual word embeddings. Furthermore, it lacks a mechanism to deal with many-to-many mappings that often show up across languages. We introduce Earth Mover's Distance to this task by providing a natural formulation that translates words in a holistic fashion, addressing the limitations of the nearest neighbor. We further extend the formulation to a new task of identifying parallel sentences, which is useful for statistical machine translation systems, thereby expanding the application realm of bilingual word embeddings. We show encouraging performance on both tasks.
Syntactic Skeleton-Based Translation
Xiao, Tong (Northeastern University) | Zhu, Jingbo (Northeastern University) | Zhang, Chunliang (Northeastern University) | Liu, Tongran (Institute of Psychology (CAS))
In this paper we propose an approach to modeling syntactically-motivated skeletal structure of source sentence for machine translation. This model allows for application of high-level syntactic transfer rules and low-level non-syntactic rules. It thus involves fully syntactic, non-syntactic, and partially syntactic derivations via a single grammar and decoding paradigm. On large-scale Chinese-English and English-Chinese translation tasks, we obtain an average improvement of +0.9 BLEU across the newswire and web genres.
Improved Neural Machine Translation with SMT Features
He, Wei (Baidu Inc.) | He, Zhongjun (Baidu Inc.) | Wu, Hua (Baidu Inc.) | Wang, Haifeng (Baidu Inc.)
Neural machine translation (NMT) conducts end-to-end translation with a source language encoder and a target language decoder, making promising translation performance. However, as a newly emerged approach, the method has some limitations. An NMT system usually has to apply a vocabulary of certain size to avoid the time-consuming training and decoding, thus it causes a serious out-of-vocabulary problem. Furthermore, the decoder lacks a mechanism to guarantee all the source words to be translated and usually favors short translations, resulting in fluent but inadequate translations. In order to solve the above problems, we incorporate statistical machine translation (SMT) features, such as a translation model and an n-gram language model, with the NMT model under the log-linear framework. Our experiments show that the proposed method significantly improves the translation quality of the state-ofthe-art NMT system on Chinese-to-English translation tasks. Our method produces a gain of up to 2.33 BLEU score on NIST open test sets.
Can machines 'learn' or 'think'? - raconteur.net
The marriage of computing power and data is finally bearing fruit in the field of cognitive computing, sometimes called machine learning or, more controversially, artificial intelligence. In its most everyday form, we see it in tools such as Google Translate or Microsoft's Bing Translate, which can translate phrases and documents effortlessly across multiple languages. More futuristically, the promise of self-driving vehicles, which can complete entire road journeys without driver intervention, is already being realised. Yet the biggest revolution in work is happening at some of the most basic levels, such as reading and dissecting legal documents to extract meaning and useful information. The tedious slog of work can be transformed by computers which are able to read and parse legal phrases, and summarise them or enter relevant details into a database or spreadsheet.