Goto

Collaborating Authors

 Machine Translation


Google introduces neural machine learning to improve translation, approach human-level accuracy The Tech Portal

#artificialintelligence

Though Google Translate is one of the most powerful language translation tools, the company still thinks there's room for major improvement. And it is now working towards creating a model which can translate phrases from one language to another automatically. Much like every other product, Google has been working on integrating machine learning translation techniques into this system as well. And today seems to be the day, we can finally see it in action. Google Neural Machine Translation system, or GNMT which utilizes state-of-the-art training techniques for improved translations has today been introduced into one of the most difficult language pair: Chinese to English.


Age of Aritificial Intelligence: How We're Already Living In a Sci-Fi Future

#artificialintelligence

When we talk about artificial intelligence (AI) most people still imagine robots who can talk, act, and behave (to a certain extent) like a human being -- like a C-3PO (Star Wars), sans the metallic look. Or maybe, a supercomputer that can read human behavior so well that it interacts seamlessly with us, while controlling the system -- like Hal 9000 (2001: A Space Odyssey) or Auto (Wall-E). While, arguably, we may not be there yet in terms of our command of AI, we are not that far. AI is definitely the direction tech development is taking, as evidenced by most recent trends, including the formation of a partnership by tech giants to push the frontier of AI. While we may not be nearing the Singularity, AI has taken leaps and bounds of improvement over the past few years alone.


Lightweight Random Indexing for Polylingual Text Classification

Journal of Artificial Intelligence Research

Multilingual Text Classification (MLTC) is a text classification task in which documents are written each in one among a set L of natural languages, and in which all documents must be classified under the same classification scheme, irrespective of language. There are two main variants of MLTC, namely Cross-Lingual Text Classification (CLTC) and Polylingual Text Classification (PLTC). In PLTC, which is the focus of this paper, we assume (differently from CLTC) that for each language in L there is a representative set of training documents; PLTC consists of improving the accuracy of each of the |L| monolingual classifiers by also leveraging the training documents written in the other (|L| โˆ’ 1) languages. The obvious solution, consisting of generating a single polylingual classifier from the juxtaposed monolingual vector spaces, is usually infeasible, since the dimensionality of the resulting vector space is roughly |L| times that of a monolingual one, and is thus often unmanageable. As a response, the use of machine translation tools or multilingual dictionaries has been proposed. However, these resources are not always available, or are not always free to use. One machine-translation-free and dictionary-free method that, to the best of our knowledge, has never been applied to PLTC before, is Random Indexing (RI). We analyse RI in terms of space and time efficiency, and propose a particular configuration of it (that we dub Lightweight Random Indexing LRI). By running experiments on two well known public benchmarks, Reuters RCV1/RCV2 (a comparable corpus) and JRC-Acquis (a parallel one), we show LRI to outperform (both in terms of effectiveness and efficiency) a number of previously proposed machine-translation-free and dictionary-free PLTC methods that we use as baselines.


A Survey of Voice Translation Methodologies - Acoustic Dialect Decoder

arXiv.org Machine Learning

Speech Translation has always been about giving source text or audio input and waiting for system to give translated output in desired form. In this paper, we present the Acoustic Dialect Decoder (ADD) - a voice to voice ear-piece translation device. We introduce and survey the recent advances made in the field of Speech Engineering, to employ in the ADD, particularly focusing on the three major processing steps of Recognition, Translation and Synthesis. We tackle the problem of machine understanding of natural language by designing a recognition unit for source audio to text, a translation unit for source language text to target language text, and a synthesis unit for target language text to target language speech. Speech from the surroundings will be recorded by the recognition unit present on the ear-piece and translation will start as soon as one sentence is successfully read. This way, we hope to give translated output as and when input is being read. The recognition unit will use Hidden Markov Models (HMMs) Based Tool-Kit (HTK), hybrid RNN systems with gated memory cells, and the synthesis unit, HMM based speech synthesis system HTS. This system will initially be built as an English to Tamil translation device.


Voice Conversion from Non-parallel Corpora Using Variational Auto-encoder

arXiv.org Machine Learning

We propose a flexible framework for spectral conversion (SC) that facilitates training with unaligned corpora. Many SC frameworks require parallel corpora, phonetic alignments, or explicit frame-wise correspondence for learning conversion functions or for synthesizing a target spectrum with the aid of alignments. However, these requirements gravely limit the scope of practical applications of SC due to scarcity or even unavailability of parallel corpora. We propose an SC framework based on variational auto-encoder which enables us to exploit non-parallel corpora. The framework comprises an encoder that learns speaker-independent phonetic representations and a decoder that learns to reconstruct the designated speaker. It removes the requirement of parallel corpora or phonetic alignments to train a spectral conversion system. We report objective and subjective evaluations to validate our proposed method and compare it to SC methods that have access to aligned corpora.


Funny Siri Responses and What it Tells us About Machine Translation โ€“ IVANNOVATION

#artificialintelligence

What do Siri and machine translation have in common? They both produce strange, sometimes ridiculous language that leave us shaking our heads with confusion. Here at IVANNOVATION we frequently use Siri as well as Google's dictation function to get our work done. Siri instantly adds items to our to do lists, adds events to our calendars, and tells us answers to important questions like, "Siri, how much wood would a woodchuck chuck if a woodchuck could chuck wood?" (Ask Siri yourself.) Likewise, Google dictation helps us avoid the ruthless onslaught of carpal tunnel syndrome by typing up our articles and emails for us.


Google Translate taps into Deep Learning to reduce errors by 60%

#artificialintelligence

The go to place for quick and easy translations โ€“ Google Translate โ€“ just received a huge upgrade with Deep Learning algorithms boosting its translation capabilities and reducing errors by 60%. Google's experiments with neural machine translation pays off in a big manner. Like most translation services, Google Translate too relied on breaking down sentences into smaller phrases or groups of words and then translated these phrases which were later joined together to produce the output. With Neural Machine Translation, Google Translate can translate entire sentences without breaking them in phrases. This new approach has been said to reduce errors by at least 60 percent compared to the previous phrase based approach.


Google Translate Gets a Deep-Learning Upgrade

#artificialintelligence

Googles engineers recently delivered a Google Translate upgrade that harnesses the popular artificial intelligence technique known as deep learning. Google has launched a Google Translate upgrade utilizing enhanced deep-learning techniques to produce more accurate translations. The neural machine translation system considers the entire sentence as one unit to be translated. The system relies on a recurrent neural network algorithm consisting of layered nodes, and a network of eight layers acts as the encoder and transforms the input into a list of vectors representing all possible meanings of each word. The second eight-layer network acts as the decoder and generates the translation one word at a time. Meanwhile, an attention network connects the encoder and decoder by directing the decoder to refer back to certain weighted vectors.


Dimension Projection among Languages based on Pseudo-relevant Documents for Query Translation

arXiv.org Artificial Intelligence

Using top-ranked documents in response to a query has been shown to be an effective approach to improve the quality of query translation in dictionary-based cross-language information retrieval. In this paper, we propose a new method for dictionary-based query translation based on dimension projection of embedded vectors from the pseudo-relevant documents in the source language to their equivalents in the target language. To this end, first we learn low-dimensional vectors of the words in the pseudo-relevant collections separately and then aim to find a query-dependent transformation matrix between the vectors of translation pairs appeared in the collections. At the next step, representation of each query term is projected to the target language and then, after using a softmax function, a query-dependent translation model is built. Finally, the model is used for query translation. Our experiments on four CLEF collections in French, Spanish, German, and Italian demonstrate that the proposed method outperforms a word embedding baseline based on bilingual shuffling and a further number of competitive baselines. The proposed method reaches up to 87% performance of machine translation (MT) in short queries and considerable improvements in verbose queries.


Google Translate Gets a Deep-Learning Upgrade

#artificialintelligence

Google Translate has become a quick-and-dirty translation solution for millions of people worldwide since it debuted a decade ago. But Google's engineers have been quietly tweaking their machine translation service's algorithms behind the scenes. They recently delivered a huge Google Translate upgrade that harnesses the popular artificial intelligence technique known as deep learning. Machine translation services such as Google Translate have mostly used a "phrase-based" approach of breaking down sentences into words and phrases to be independently translated. But several years ago, Google began experimenting with a deep-learning technique, called neural machine translation, that can translate entire sentences without breaking them down into smaller components.