Setting up the Data Printer with Improved English to Ukrainian Machine Translation

Paniv, Yurii, Chaplynskyi, Dmytro, Trynus, Nikita, Kyrylov, Volodymyr

Jul-12-2024–arXiv.org Artificial Intelligence

To build large language models for Ukrainian we need to expand our corpora with large amounts of new algorithmic tasks expressed in natural language. Examples of task performance expressed in English are abundant, so with a high-quality translation system our community will be enabled to curate datasets faster. To aid this goal, we introduce a recipe to build a translation system using supervised finetuning of a large pretrained language model with a noisy parallel dataset of 3M pairs of Ukrainian and English sentences followed by a second phase of training using 17K examples selected by k-fold perplexity filtering on another dataset of higher quality. Our decoder-only model named Dragoman beats performance of previous state of the art encoder-decoder models on the FLORES devtest set.

large language model, natural language, translation, (16 more...)

arXiv.org Artificial Intelligence

Jul-12-2024

arXiv.org PDF

Add feedback

Country:
- Indian Ocean (0.04)
- Africa (0.04)
- Oceania > Australia
  - New South Wales (0.04)
- North America
  - Dominican Republic (0.04)
  - United States
    - Washington > King County
      - Seattle (0.04)
    - Pennsylvania > Philadelphia County
      - Philadelphia (0.04)
    - Massachusetts > Suffolk County
      - Boston (0.04)
  - Canada > Ontario
    - Toronto (0.04)
- Europe
  - Ukraine > Kyiv Oblast
    - Kyiv (0.04)
  - Portugal > Lisbon
    - Lisbon (0.04)
  - Italy > Tuscany
    - Florence (0.04)
  - Finland > Pirkanmaa
    - Tampere (0.04)
  - Croatia > Dubrovnik-Neretva County
    - Dubrovnik (0.04)
  - Belgium > Brussels-Capital Region
    - Brussels (0.04)
- Asia > Middle East
  - UAE > Abu Dhabi Emirate > Abu Dhabi (0.05)

Genre:
- Research Report (0.82)

Technology:
- Information Technology > Artificial Intelligence > Natural Language
  - Machine Translation (1.00)
  - Large Language Model (1.00)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found