Goto

Collaborating Authors

 Machine Translation


WIPO Develops Cutting-Edge Translation Tool For Patent Documents

#artificialintelligence

The World Intellectual Property Organization has developed a ground-breaking new "artificial intelligence"-based translation tool for patent documents, handing innovators around the world the highest-quality service yet available for accessing information on new technologies. WIPO Translate now incorporates cutting-edge neural machine translation technology to render highly technical patent documents into a second language in a style and syntax that more closely mirrors common usage, out-performing other translation tools built on previous technologies. WIPO has initially "trained" the new technology to translate Chinese, Japanese and Korean patent documents into English. Patent applications in those languages accounted for some 55% of worldwide filings in 20141. Users can already try out the Chinese-English translation facility on the public beta test platform.


The Limits of Modern AI: A Story The Best Schools

#artificialintelligence

The dream of thinking machines goes back centuries, at least to Gottfried Wilhelm Leibniz, in the 17th century. Leibniz (right) helped invent mechanical calculators, independently of Isaac Newton developed the integral calculus, and had a lifelong fascination with reducing thinking to calculation. His Mathesis Universalis was a vision of universal science made possible by a mathematical language more precise than natural languages, like English. The Limits of Modern AI: A Story In the 18th Century the Enlightenment philosopher and proto-psychologist ร‰tienne Bonnot de Condillac imagined a statue outwardly appearing like a man and also with what he called "the inward organization." In an example of supreme armchair speculation, Condillac imagined pouring facts--bits of knowledge--into its head, wondering when intelligence would emerge. Condillac's musings drew inspiration from the early mechanical philosophy of Thomas Hobbes, who had famously declared that thinking was nothing but ...


Google introduces neural machine learning to improve translation, approach human-level accuracy The Tech Portal

#artificialintelligence

Though Google Translate is one of the most powerful language translation tools, the company still thinks there's room for major improvement. And it is now working towards creating a model which can translate phrases from one language to another automatically. Much like every other product, Google has been working on integrating machine learning translation techniques into this system as well. And today seems to be the day, we can finally see it in action. Google Neural Machine Translation system, or GNMT which utilizes state-of-the-art training techniques for improved translations has today been introduced into one of the most difficult language pair: Chinese to English.


Age of Aritificial Intelligence: How We're Already Living In a Sci-Fi Future

#artificialintelligence

When we talk about artificial intelligence (AI) most people still imagine robots who can talk, act, and behave (to a certain extent) like a human being -- like a C-3PO (Star Wars), sans the metallic look. Or maybe, a supercomputer that can read human behavior so well that it interacts seamlessly with us, while controlling the system -- like Hal 9000 (2001: A Space Odyssey) or Auto (Wall-E). While, arguably, we may not be there yet in terms of our command of AI, we are not that far. AI is definitely the direction tech development is taking, as evidenced by most recent trends, including the formation of a partnership by tech giants to push the frontier of AI. While we may not be nearing the Singularity, AI has taken leaps and bounds of improvement over the past few years alone.


Lightweight Random Indexing for Polylingual Text Classification

Journal of Artificial Intelligence Research

Multilingual Text Classification (MLTC) is a text classification task in which documents are written each in one among a set L of natural languages, and in which all documents must be classified under the same classification scheme, irrespective of language. There are two main variants of MLTC, namely Cross-Lingual Text Classification (CLTC) and Polylingual Text Classification (PLTC). In PLTC, which is the focus of this paper, we assume (differently from CLTC) that for each language in L there is a representative set of training documents; PLTC consists of improving the accuracy of each of the |L| monolingual classifiers by also leveraging the training documents written in the other (|L| โˆ’ 1) languages. The obvious solution, consisting of generating a single polylingual classifier from the juxtaposed monolingual vector spaces, is usually infeasible, since the dimensionality of the resulting vector space is roughly |L| times that of a monolingual one, and is thus often unmanageable. As a response, the use of machine translation tools or multilingual dictionaries has been proposed. However, these resources are not always available, or are not always free to use. One machine-translation-free and dictionary-free method that, to the best of our knowledge, has never been applied to PLTC before, is Random Indexing (RI). We analyse RI in terms of space and time efficiency, and propose a particular configuration of it (that we dub Lightweight Random Indexing LRI). By running experiments on two well known public benchmarks, Reuters RCV1/RCV2 (a comparable corpus) and JRC-Acquis (a parallel one), we show LRI to outperform (both in terms of effectiveness and efficiency) a number of previously proposed machine-translation-free and dictionary-free PLTC methods that we use as baselines.


Voice Conversion from Non-parallel Corpora Using Variational Auto-encoder

arXiv.org Machine Learning

We propose a flexible framework for spectral conversion (SC) that facilitates training with unaligned corpora. Many SC frameworks require parallel corpora, phonetic alignments, or explicit frame-wise correspondence for learning conversion functions or for synthesizing a target spectrum with the aid of alignments. However, these requirements gravely limit the scope of practical applications of SC due to scarcity or even unavailability of parallel corpora. We propose an SC framework based on variational auto-encoder which enables us to exploit non-parallel corpora. The framework comprises an encoder that learns speaker-independent phonetic representations and a decoder that learns to reconstruct the designated speaker. It removes the requirement of parallel corpora or phonetic alignments to train a spectral conversion system. We report objective and subjective evaluations to validate our proposed method and compare it to SC methods that have access to aligned corpora.


A Survey of Voice Translation Methodologies - Acoustic Dialect Decoder

arXiv.org Machine Learning

Speech Translation has always been about giving source text or audio input and waiting for system to give translated output in desired form. In this paper, we present the Acoustic Dialect Decoder (ADD) - a voice to voice ear-piece translation device. We introduce and survey the recent advances made in the field of Speech Engineering, to employ in the ADD, particularly focusing on the three major processing steps of Recognition, Translation and Synthesis. We tackle the problem of machine understanding of natural language by designing a recognition unit for source audio to text, a translation unit for source language text to target language text, and a synthesis unit for target language text to target language speech. Speech from the surroundings will be recorded by the recognition unit present on the ear-piece and translation will start as soon as one sentence is successfully read. This way, we hope to give translated output as and when input is being read. The recognition unit will use Hidden Markov Models (HMMs) Based Tool-Kit (HTK), hybrid RNN systems with gated memory cells, and the synthesis unit, HMM based speech synthesis system HTS. This system will initially be built as an English to Tamil translation device.


Funny Siri Responses and What it Tells us About Machine Translation โ€“ IVANNOVATION

#artificialintelligence

What do Siri and machine translation have in common? They both produce strange, sometimes ridiculous language that leave us shaking our heads with confusion. Here at IVANNOVATION we frequently use Siri as well as Google's dictation function to get our work done. Siri instantly adds items to our to do lists, adds events to our calendars, and tells us answers to important questions like, "Siri, how much wood would a woodchuck chuck if a woodchuck could chuck wood?" (Ask Siri yourself.) Likewise, Google dictation helps us avoid the ruthless onslaught of carpal tunnel syndrome by typing up our articles and emails for us.


Google Translate taps into Deep Learning to reduce errors by 60%

#artificialintelligence

The go to place for quick and easy translations โ€“ Google Translate โ€“ just received a huge upgrade with Deep Learning algorithms boosting its translation capabilities and reducing errors by 60%. Google's experiments with neural machine translation pays off in a big manner. Like most translation services, Google Translate too relied on breaking down sentences into smaller phrases or groups of words and then translated these phrases which were later joined together to produce the output. With Neural Machine Translation, Google Translate can translate entire sentences without breaking them in phrases. This new approach has been said to reduce errors by at least 60 percent compared to the previous phrase based approach.


Google Translate Gets a Deep-Learning Upgrade

#artificialintelligence

Googles engineers recently delivered a Google Translate upgrade that harnesses the popular artificial intelligence technique known as deep learning. Google has launched a Google Translate upgrade utilizing enhanced deep-learning techniques to produce more accurate translations. The neural machine translation system considers the entire sentence as one unit to be translated. The system relies on a recurrent neural network algorithm consisting of layered nodes, and a network of eight layers acts as the encoder and transforms the input into a list of vectors representing all possible meanings of each word. The second eight-layer network acts as the decoder and generates the translation one word at a time. Meanwhile, an attention network connects the encoder and decoder by directing the decoder to refer back to certain weighted vectors.