Authorship Attribution in Multilingual Machine-Generated Texts
La Cava, Lucio, Macko, Dominik, Móro, Róbert, Srba, Ivan, Tagarelli, Andrea
–arXiv.org Artificial Intelligence
As Large Language Models (LLMs) have reached human-like fl uency and coherence, distinguishing machine-generated text (MGT) from human-written content becomes increasingly dif fi cult. While early efforts in MGT detection have focused on binary classi fi cation, the growing landscape and diversity of LLMs require a more fi ne-grained yet challenging authorship attribution (AA), i.e., being able to identify the precise generator (LLM or human) behind a text. However, AA remains nowadays con fi ned to a monolingual setting, with English being the most investigated one, overlooking the multilingual nature and usage of modern LLMs. In this work, we introduce the problem of Multilingual Authorship Attribution, which involves attributing texts to human or multiple LLM generators across diverse languages. Focusing on 18 languages--covering multiple families and writing scripts--and 8 generators (7 LLMs and the human-authored class), we investigate the multilingual suitability of monolingual AA methods, their cross-lingual transferability, and the impact of generators on attribution performance. Our results reveal that while certain monolingual AA methods can be adapted to multilingual settings, signi fi cant limitations and challenges remain, particularly in transferring across diverse language families, underscoring the complexity of multilingual AA and the need for more robust approaches to better match real-world scenarios.
arXiv.org Artificial Intelligence
Aug-5-2025
- Country:
- Europe > Austria (0.28)
- North America
- United States (0.28)
- Mexico (0.28)
- Genre:
- Research Report > New Finding (0.48)
- Industry:
- Information Technology > Security & Privacy (0.46)
- Technology: