Goto

Collaborating Authors

 Large Language Model


ProPILE: Probing Privacy Leakage in Large Language Models Siwon Kim 1, Sangdoo Y un 3 Hwaran Lee 3 Martin Gubri

Neural Information Processing Systems

The rapid advancement and widespread use of large language models (LLMs) have raised significant concerns regarding the potential leakage of personally identifiable information (PII). These models are often trained on vast quantities of web-collected data, which may inadvertently include sensitive personal data.








Appendix T able of Contents

Neural Information Processing Systems

For all baseline methods, we use the MinMaxScaler from sklearn. The likelihood of generating the validation conditioned on the remaining training series is used to select the hyperparameters. We compare the performance of our GPT -3 predictor against popular time series models. GPT -3 continues to be competitive with or outperforms the baselines on all of the tasks, from in-context learning alone. GPT -3's performance is not due to memorization of the test data. Even if our evaluation datasets are present in the GPT -3 training data, it's unlikely that GPT -3's good performance is the result of memorization for at least two reasons a priori.