Accuracy
Watermarking Makes Language Models Radioactive Tom Sander
Current methods like membership inference or active IP protection either work only in settings where the suspected text is known or do not provide reliable statistical guarantees. We discover that, on the contrary, it is possible to reliably determine if a language model was trained on synthetic data if that data is output by a watermarked LLM.
RankFeat: Rank-1 Feature Removal for Out-of-distribution Detection-Supplementary Material-A Experimental Setup Implementation Details
Table 1 present the evaluation results. We also evaluate our method on the CIFAR benchmark with various model architectures. The best three results are highlighted with red, blue, and cyan. The best two results are highlighted with red and blue . Table 2 compares the performance against all the post hoc baselines.