AutoPureData: Automated Filtering of Web Data for LLM Fine-tuning
–arXiv.org Artificial Intelligence
Keeping models up-todate is crucial in domains where data changes frequently, such as news and academic research. Unfortunately, very few LLMs are continuously updated, as they do not integrate the latest data. Using search engines on demand is often time-consuming and expensive, and web data is reliable on proper filtration. This research focuses on regular automated web data collection and filtration to support up-to-date Responsible AI models. AI safety is crucial for the success of Responsible AI models. Data used for training Responsible AI models should be both safe and unbiased. As "garbage in, garbage out" suggests, the input data for training or fine-tuning an LLM impacts the quality of the model [1]. The quality of the model depends on the quality of the data used to train or fine-tune it. The web is a vast source of information, but its reliability varies significantly.
arXiv.org Artificial Intelligence
Jun-27-2024