Tagengo: A Multilingual Chat Dataset

Devine, Peter

arXiv.org Artificial Intelligence 

Recently, open source large language models We found that our model achieved better evaluation (LLMs) have grown drastically in both popularity scores on multilingual chat benchmarks compared and performance. Models such as Llama to the similarly sized state-of-the-art open 3 (AI@Meta, 2024b) have exceeded the performance source models, indicating the high quality and of previous state-of-the-art proprietary models diversity of our training dataset. We also find like GPT3.5 (Ouyang et al., 2022) on popular that our multilingual-trained LLM performs better robust benchmarks including the Chatbot Arena on Japanese chat benchmarks compared to our leaderboard (Chiang et al., 2024). These open Japanese-only-trained LLM, indicating that transfer source LLMs are also increasingly being used in learning from training on other languages is commerical AI chat products such as the Meta AI beneficial for training even monolingual models assistant (AI@Meta, 2024a).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found