How Good Are Synthetic Requirements ? Evaluating LLM-Generated Datasets for AI4RE

El-Hajjami, Abdelkarim, Salinesi, Camille

arXiv.org Artificial Intelligence 

The shortage of publicly available, labeled requirements datasets represents a major barrier to advancing Artificial Intelligence for Requirements Engineering (AI4RE) research. While Large Language Models offer promising capabilities for synthetic data generation, systematic approaches for controlling and optimizing the quality of generated requirements remain underexplored. This paper presents Synthline v1, an enhanced Product Line approach for generating synthetic requirements data that extends our earlier v0 version with advanced generation strategies and curation techniques. We investigate four research questions examining how prompting strategies, automated prompt optimization, and post-generation curation affect synthetic data quality across four requirements classification tasks: defects detection, functional vs. nonfunctional classification, quality vs. non-quality classification, and security vs. non-security classification. Our experimental evaluation demonstrates that multi-sample prompting significantly improves both utility and diversity compared to single-sample generation, with F1-score improvements ranging from 6-44 percentage points. The application of PACE (Prompt Actor-Critic Editing) for automated prompt optimization shows task-dependent results, substantially improving functional requirements classification (+32.5 percentage points) while degrading performance on other tasks. Counterintuitively, similarity-based curation consistently improves diversity metrics but often hurts classification performance, revealing that not all redundancy is detrimental to Machine Learning effectiveness. Most significantly, our results establish that synthetic requirements can match or surpass human-authored requirements for specific classification tasks, with synthetic data achieving superior performance for security (+7.8 percentage points) and defects classification (+15.4 percentage points). These findings provide actionable insights for the AI4RE community and demonstrate a viable path toward addressing dataset scarcity through systematic synthetic data generation.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found