Corpus and Models for Lemmatisation and POS-tagging of Classical French Theatre
Camps, Jean-Baptiste, Gabay, Simon, Fièvre, Paul, Clérice, Thibault, Cafiero, Florian
–arXiv.org Artificial Intelligence
This paper describes the process of building an annotated corpus and training models for classical French literature, with a focus on theatre, and particularly comedies in verse. It was originally developed as a preliminary step to the stylometric analyses presented in Cafiero and Camps [2019]. The use of a recent lemmatiser based on neural networks and a CRF tagger allows to achieve accuracies beyond the current state-of-the art on the in-domain test, and proves to be robust during out-of-domain tests, i.e. up to 20th c. novels. I INTRODUCTION If many lemmatisers and POS taggers have been trained, and sometimes conceived, for French (e.g. Tellier et al. [2012], Urieli [2013]...), they usually focus on contemporary French and tools for Ancien Régime French remain scarce. One important exception is the TreeTagger [Schmid, 1995] model developed by Diwersy et al. [2017] for the Presto project [Vigier and Blumenthal, 2013-2017].
arXiv.org Artificial Intelligence
Feb-5-2021
- Country:
- Asia (0.04)
- North America > Costa Rica
- Heredia Province > Heredia (0.04)
- Europe
- Spain > Aragón (0.04)
- Switzerland > Neuchâtel
- Neuchâtel (0.04)
- Portugal > Lisbon
- Lisbon (0.04)
- Germany
- Saxony > Leipzig (0.04)
- Baden-Württemberg > Stuttgart Region
- Stuttgart (0.04)
- France
- Occitanie > Haute-Garonne
- Toulouse (0.04)
- Grand Est > Meurthe-et-Moselle
- Nancy (0.04)
- Occitanie > Haute-Garonne
- Genre:
- Research Report (1.00)
- Technology: