Open Korean Corpora: A Practical Report

Cho, Won Ik, Moon, Sangwhan, Song, Youngsook

May-16-2023–arXiv.org Artificial Intelligence

Korean is often referred to as a low-resource language in the research community. While this claim is partially true, it is also because the availability of resources is inadequately advertised and curated. This work curates and reviews a list of Korean corpora, first describing institution-level resource development, then further iterate through a list of current open datasets for different types of tasks. We then propose a direction on how open-source dataset construction and releases should be done for less-resourced languages to promote research.

large language model, machine learning, natural language, (17 more...)

arXiv.org Artificial Intelligence

May-16-2023

arXiv.org PDF

Add feedback

Country:
- North America > United States
  - Oregon (0.04)
  - Washington > King County
    - Seattle (0.04)
  - California > San Diego County
    - San Diego (0.04)
- Europe
  - France > Provence-Alpes-Côte d'Azur
    - Bouches-du-Rhône > Marseille (0.04)
  - Croatia > Dubrovnik-Neretva County
    - Dubrovnik (0.04)
- Asia
  - North Korea (0.04)
  - South Korea
    - Seoul > Seoul (0.05)
    - Gyeongsangnam-do > Changwon (0.04)
  - Middle East > UAE
    - Abu Dhabi Emirate > Abu Dhabi (0.04)
  - Japan > Honshū
    - Kantō > Tokyo Metropolis Prefecture > Tokyo (0.14)

Genre:
- Research Report (0.40)

Technology:
- Information Technology
  - Communications > Social Media (1.00)
  - Artificial Intelligence
    - Speech (1.00)
    - Machine Learning (0.93)
    - Natural Language
      - Text Processing (0.94)
      - Large Language Model (0.93)
      - Grammars & Parsing (0.69)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found