A Multi-dimensional Evaluation of Tokenizer-free Multilingual Pretrained Models

Sun, Jimin, Fernandes, Patrick, Wang, Xinyi, Neubig, Graham

arXiv.org Artificial Intelligence 

In a multilingual setting, subword tokenization Several recent results (Clark et al., 2022; Xue can be sub-optimal as supporting hundreds et al., 2022) have excited the research community of languages with various scripts and vocabulary with the possibility of "tokenizer-free" models, causes segmentation mismatch between languages character-level and byte-level models, as an alternative and over-segmentation in the lower-resourced languages to more traditional subword-based models. We, (Wang et al., 2020; Ebrahimi and Kann, the authors of this paper, were also initially excited 2021). To alleviate this problem, recent works by these results - the possibility of eschewing the propose to remove the preprocessing step of subword two-step processing pipeline of subword segmentation segmentation by directly using characters or and subword-based models would reduce the bytes as lexical units (Clark et al., 2022; Xue et al., corresponding difficulties in cross-lingual transfer 2022). Tab. 1 presents an overview of the different (Hu et al., 2020; Maronikolakis et al., 2021; Rust tokenizer-free multilingual models with comparable et al., 2021; Wang et al., 2021) or domain adaptation subword models. Next, we briefly describe (Sato et al., 2020; Liu et al., 2021) due to inconsistent the two tokenizer-free models we consider in this subword units. However, upon several attempts work.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found