LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
Cegin, Jan, Simko, Jakub, Brusilovsky, Peter
–arXiv.org Artificial Intelligence
The generative large language models (LLMs) are increasingly being used for data augmentation tasks, where text samples are LLM-paraphrased and then used for classifier fine-tuning. However, a research that would confirm a clear cost-benefit advantage of LLMs over more established augmentation methods is largely missing. To study if (and when) is the LLM-based augmentation advantageous, we compared the effects of recent LLM augmentation methods with established ones on 6 datasets, 3 classifiers and 2 fine-tuning methods. We also varied the number of seeds and collected samples to better explore the downstream model accuracy space. Finally, we performed a cost-benefit analysis and show that LLM-based methods are worthy of deployment only when very small number of seeds is used. Moreover, in many cases, established methods lead to similar or better model accuracies.
arXiv.org Artificial Intelligence
Aug-29-2024
- Country:
- North America
- Dominican Republic (0.04)
- United States
- Pennsylvania (0.04)
- Washington > King County
- Seattle (0.04)
- New York > New York County
- New York City (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Canada > Ontario
- Toronto (0.04)
- Europe
- Germany > Berlin (0.04)
- Spain > Catalonia
- Barcelona Province > Barcelona (0.04)
- Slovakia > Bratislava
- Bratislava (0.04)
- Czechia > South Moravian Region
- Brno (0.04)
- Asia
- North America
- Genre:
- Research Report
- New Finding (1.00)
- Experimental Study (0.68)
- Research Report
- Technology: