Unsupervised Neural Machine Translation with Generative Language Models Only

Han, Jesse Michael, Babuschkin, Igor, Edwards, Harrison, Neelakantan, Arvind, Xu, Tao, Polu, Stanislas, Ray, Alex, Shyam, Pranav, Ramesh, Aditya, Radford, Alec, Sutskever, Ilya

Oct-11-2021–arXiv.org Artificial Intelligence

We show how to derive state-of-the-art unsupervised neural machine translation systems from generatively pre-trained language models. Our method consists of three steps: few-shot amplification, distillation, and backtranslation. We first use the zero-shot translation ability of large pre-trained language models to generate translations for a small set of unlabeled sentences. We then amplify these zero-shot translations by using them as few-shot demonstrations for sampling a larger synthetic dataset. This dataset is distilled by discarding the few-shot demonstrations and then fine-tuning. During backtranslation, we repeatedly generate translations for a set of inputs and then fine-tune a single language model on both directions of the translation task at once, ensuring cycle-consistency by swapping the roles of gold monotext and generated translations when fine-tuning. By using our method to leverage GPT-3's zero-shot translation capability, we achieve a new state-of-the-art in unsupervised translation on the WMT14 English-French benchmark, attaining a BLEU score of 42.1.

artificial intelligence, machine translation, natural language, (15 more...)

arXiv.org Artificial Intelligence

Oct-11-2021

arXiv.org PDF

Add feedback

Country:
- Asia (0.68)
- Europe (1.00)
- North America
  - Canada > British Columbia
    - Metro Vancouver Regional District > Vancouver (0.14)
  - United States
    - California > Los Angeles County
      - Long Beach (0.14)
    - Minnesota > Hennepin County
      - Minneapolis (0.14)

Genre:
- Research Report (0.52)

Technology:
- Information Technology > Artificial Intelligence
  - Machine Learning > Neural Networks
    - Deep Learning (0.91)
  - Natural Language > Machine Translation (1.00)