Deep Learning
ProPILE: Probing Privacy Leakage in Large Language Models Siwon Kim 1, Sangdoo Y un 3 Hwaran Lee 3 Martin Gubri
The rapid advancement and widespread use of large language models (LLMs) have raised significant concerns regarding the potential leakage of personally identifiable information (PII). These models are often trained on vast quantities of web-collected data, which may inadvertently include sensitive personal data.
Appendix: Combating Representation Learning Disparity with Geometric Harmonization
We provide our source codes to ensure the reproducibility of our experimental results. Below we summarize several critical aspects w.r .tthe The datasets we used are all publicly accessible, which is introduced in Appendix E.1. For long-tailed subsets, we strictly follows previous work [29] on CIFAR-100-L T to avoid the bias attribute to the sampling randomness. On ImageNet-L T and Places-L T, we employ the widely-used data split first introduced in [44]. All the experiments are conducted on NVIDIA GeForce RTX 3090 with Python 3.7 and Pytorch 1.7.