Iterative Learning of Parallel Lexicons and Phrases from Non-Parallel Corpora

Jul-15-2015–AAAI Conferences

While parallel corpora are an indispensable resource for data-driven multilingual natural language processing tasks such as machine translation, they are limited in quantity, quality and coverage. As a result, learning translation models from non-parallel corpora has become increasingly important nowadays, especially for low-resource languages. In this work, we propose a joint model for iteratively learning parallel lexicons and phrases from nonparallel corpora. The model is trained using a Viterbi EM algorithm that alternates between constructing parallel phrases using lexicons and updating lexicons based on the constructed parallel phrases. Experiments on Chinese-English datasets show that our approach learns better parallel lexicons and phrases and improves translation performance significantly.

corpora, english phrase, translation, (15 more...)

AAAI Conferences

Jul-15-2015

Conferences PDF

Add feedback

Country:
- Asia
  - Singapore (0.04)
  - China
    - Jiangsu Province (0.04)
    - Beijing > Beijing (0.04)

Technology:
- Information Technology > Artificial Intelligence > Natural Language > Machine Translation (1.00)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found