A Novel Dataset for Financial Education Text Simplification in Spanish

Perez-Rojas, Nelson, Calderon-Ramirez, Saul, Solis-Salazar, Martin, Romero-Sandoval, Mario, Arias-Monge, Monica, Saggion, Horacio

Dec-15-2023–arXiv.org Artificial Intelligence

Text simplification, crucial in natural language processing, aims to make texts more comprehensible, particularly for specific groups like visually impaired Spanish speakers, a less-represented language in this field. In Spanish, there are few datasets that can be used to create text simplification systems. Our research has the primary objective to develop a Spanish financial text simplification dataset. We created a dataset with 5,314 complex and simplified sentence pairs using established simplification rules. We also compared our dataset with the simplifications generated from GPT-3, Tuner, and MT5, in order to evaluate the feasibility of data augmentation using these systems. In this manuscript we present the characteristics of our dataset and the findings of the comparisons with other systems. The dataset is available at Hugging face, saul1917/FEINA.

dataset, simplification, text segment, (15 more...)

arXiv.org Artificial Intelligence

Dec-15-2023

arXiv.org PDF

Add feedback

Country:
- South America > Ecuador
  - Guayas Province > Guayaquil (0.04)
- North America
  - Costa Rica (0.05)
  - United States (0.04)
- Europe
  - Spain > Galicia
    - Madrid (0.04)
  - France > Provence-Alpes-Côte d'Azur
    - Bouches-du-Rhône > Marseille (0.04)
  - Denmark > Capital Region
    - Copenhagen (0.04)
- Asia > China
  - Hong Kong (0.04)

Genre:
- Research Report
  - New Finding (1.00)
  - Experimental Study (0.68)

Industry:
- Health & Medicine (1.00)

Technology:
- Information Technology > Artificial Intelligence
  - Natural Language > Large Language Model (1.00)
  - Machine Learning > Neural Networks
    - Deep Learning (1.00)