A Multi-dimensional Evaluation of Tokenizer-free Multilingual Pretrained Models
Sun, Jimin, Fernandes, Patrick, Wang, Xinyi, Neubig, Graham
–arXiv.org Artificial Intelligence
In a multilingual setting, subword tokenization Several recent results (Clark et al., 2022; Xue can be sub-optimal as supporting hundreds et al., 2022) have excited the research community of languages with various scripts and vocabulary with the possibility of "tokenizer-free" models, causes segmentation mismatch between languages character-level and byte-level models, as an alternative and over-segmentation in the lower-resourced languages to more traditional subword-based models. We, (Wang et al., 2020; Ebrahimi and Kann, the authors of this paper, were also initially excited 2021). To alleviate this problem, recent works by these results - the possibility of eschewing the propose to remove the preprocessing step of subword two-step processing pipeline of subword segmentation segmentation by directly using characters or and subword-based models would reduce the bytes as lexical units (Clark et al., 2022; Xue et al., corresponding difficulties in cross-lingual transfer 2022). Tab. 1 presents an overview of the different (Hu et al., 2020; Maronikolakis et al., 2021; Rust tokenizer-free multilingual models with comparable et al., 2021; Wang et al., 2021) or domain adaptation subword models. Next, we briefly describe (Sato et al., 2020; Liu et al., 2021) due to inconsistent the two tokenizer-free models we consider in this subword units. However, upon several attempts work.
arXiv.org Artificial Intelligence
Oct-13-2022
- Country:
- Oceania > Australia (0.04)
- North America
- Europe
- Genre:
- Research Report > New Finding (0.46)
- Technology: