Machine Translation
Artificial intelligence goes bilingual--without a dictionary
Computers might soon translate between many more languages. Automatic language translation has come a long way, thanks to neural networks--computer algorithms that take inspiration from the human brain. But training such networks requires an enormous amount of data: millions of sentence-by-sentence translations to demonstrate how a human would do it. Now, two new papers show that neural networks can learn to translate with no parallel texts--a surprising advance that could make documents in many languages more accessible. "Imagine that you give one person lots of Chinese books and lots of Arabic books--none of them overlapping--and the person has to learn to translate Chinese to Arabic. That seems impossible, right?" says the first author of one study, Mikel Artetxe, a computer scientist at the University of the Basque Country (UPV) in San Sebastiร n, Spain.
AI's sharing economy: Microsoft creates publicly available datasets
Samira Ebrahimi Kahou and her colleagues at Microsoft Research Maluuba recently set out to solve an interesting research problem: How could they use artificial intelligence to correctly reason about information found in graphs and pie charts? One big obstacle, they discovered, was that the research area was so new that there weren't any existing datasets available for them to test their hypotheses. The FigureQA dataset, which the team released publicly earlier this fall, is one of a number of datasets, metrics and other tools for testing AI systems that Microsoft researchers and engineers have created and shared in recent years. Researchers all over the world use them to see how well their AI systems do at everything from translating conversational speech to predicting the next word a person may want to type. The teams say these tools provide a codified way for everyone from academic researchers to industry experts to test their systems, compare their work and learn from each other.
Investors Shovel Millions into Natural Language Processing Slator
Among the vast business applications of artificial intelligence, Slator has been keeping a close eye on neural machine translation (MT). However, the boundaries between MT and broader tech like natural language processing (NLP) are sometimes fuzzy. The services resulting from these technologies are often adjacent: translation on one side and chatbots on another. In fact, some companies combine them into a single service--multilingual chatbots, for instance. This is why a recent slew of significant funding rounds in the NLP space has caught our attention. In June 2017, Italy-based venture incubator H-Farm acquired language technology services provider CELI, also headquartered in Italy, in a leveraged buyout of 100% of its shares.
Word embeddings in 2017: Trends and future directions
The word2vec method based on skip-gram with negative sampling (Mikolov et al., 2013) [49] was published in 2013 and had a large impact on the field, mainly through its accompanying software package, which enabled efficient training of dense word representations and a straightforward integration into downstream models. In some respects, we have come far since then: Word embeddings have established themselves as an integral part of Natural Language Processing (NLP) models. In other aspects, we might as well be in 2013 as we have not found ways to pre-train word embeddings that have managed to supersede the original word2vec. This post will focus on the deficiencies of word embeddings and how recent approaches have tried to resolve them. If not otherwise stated, this post discusses pre-trained word embeddings, i.e. word representations that have been learned on a large corpus using word2vec and its variants.
Artificial Intelligence Is Now Your Coworker
Last fall, Google Translate rolled out a new-and-improved artificial intelligence translation engine that it claimed was, at times, "nearly indistinguishable" from human translation. Jost Zetzsche could only roll his eyes. The German native had been working as a professional translator for 20 years, and he'd heard time and time again that his industry would be threatened by advances in automation. Every time, he'd found, the hype was overblown--and Google Translate's makeover was no exception. It certainly wasn't the key to translation, he thought.
A Survey on Lexical Simplification
Paetzold, Gustavo H., Specia, Lucia
Lexical Simplification is the process of replacing complex words in a given sentence with simpler alternatives of equivalent meaning. This task has wide applicability both as an assistive technology for readers with cognitive impairments or disabilities, such as Dyslexia and Aphasia, and as a pre-processing tool for other Natural Language Processing tasks, such as machine translation and summarisation. The problem is commonly framed as a pipeline of four steps: the identification of complex words, the generation of substitution candidates, the selection of those candidates that fit the context, and the ranking of the selected substitutes according to their simplicity. In this survey we review the literature for each step in this typical Lexical Simplification pipeline and provide a benchmarking of existing approaches for these steps on publicly available datasets. We also provide pointers for datasets and resources available for the task.
"Found in Translation": Predicting Outcomes of Complex Organic Chemistry Reactions using Neural Sequence-to-Sequence Models
Schwaller, Philippe, Gaudin, Theophile, Lanyi, David, Bekas, Costas, Laino, Teodoro
There is an intuitive analogy of an organic chemist's understanding of a compound and a language speaker's understanding of a word. Consequently, it is possible to introduce the basic concepts and analyze potential impacts of linguistic analysis to the world of organic chemistry. In this work, we cast the reaction prediction task as a translation problem by introducing a template-free sequence-to-sequence model, trained end-to-end and fully data-driven. We propose a novel way of tokenization, which is arbitrarily extensible with reaction information. With this approach, we demonstrate results superior to the state-of-the-art solution by a significant margin on the top-1 accuracy. Specifically, our approach achieves an accuracy of 80.1% without relying on auxiliary knowledge such as reaction templates. Also, 66.4% accuracy is reached on a larger and noisier dataset.
Microsoft Monday: Free Windows 10 Upgrade Fully Ending, Xbox One X Events, Sync Xbox Settings
"Microsoft Monday" is a weekly column that focuses on all things Microsoft. This week, Microsoft Monday includes details about the Windows 10 upgrade program fully ending, the shutdown of Outlook.com When Microsoft originally released Windows 10, it was available as a free upgrade PCs running on Windows 7 and Windows 8.1. After a year, the free upgrade offer was extended for people that were seeking enhanced accessibility features. This workaround will no longer be available after December 31, 2017.