PL-MTEB: Polish Massive Text Embedding Benchmark

Poświata, Rafał, Dadas, Sławomir, Perełkiewicz, Michał

arXiv.org Artificial Intelligence 

In many cases, they are fundamental elements of the created systems and significantly impact their performance. Therefore, it is important to select the appropriate embedding model based on the results of its evaluation. Most often, evaluation is conducted on individual tasks using a limited set of datasets, leaving the open question of how such embedding models would work for other tasks. To solve this problem, Muennighoff et al. [2023] created a Massive Text Embedding Benchmark (MTEB). MTEB provides a simple and clear way to examine how the model behaves for different types of tasks. Most of the tasks in MTEB are based on English-language datasets, and only a few are multilingual, making it impossible to do a good comparison of models for languages other than English. Therefore, extensions to MTEB with language-specific task sets have begun to appear, among which are C-MTEB [Xiao et al., 2023] for Chinese, F-MTEB

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found