Hakim: Farsi Text Embedding Model
Sarmadi, Mehran, Alikhani, Morteza, Zinvandi, Erfan, Pourbahman, Zahra
–arXiv.org Artificial Intelligence
Recent advancements in text embedding have significantly improved natural language understanding across many languages, yet Persian remains notably underrepresented in large-scale embedding research. In this paper, we present Hakim, a novel state-of-the-art Persian text embedding model that achieves a 8.5% performance improvement over existing approaches on the FaMTEB benchmark, outperforming all previously developed Persian language models. As part of this work, we introduce three new datasets - Corpesia, Pairsia-sup, and Pairsia-unsup - to support supervised and unsupervised training scenarios. Additionally, Hakim is designed for applications in chatbots and retrieval-augmented generation (RAG) systems, particularly addressing retrieval tasks that require incorporating message history within these systems. We also propose a new baseline model built on the BERT architecture. Our language model consistently achieves higher accuracy across various Persian NLP tasks, while the RetroMAE-based model proves particularly effective for textual information retrieval applications. Together, these contributions establish a new foundation for advancing Persian language understanding.
arXiv.org Artificial Intelligence
Oct-10-2025
- Country:
- North America > United States (0.28)
- Asia > Middle East (0.28)
- Genre:
- Research Report (0.64)
- Technology: