Goto

Collaborating Authors

 Accuracy




Watermarking Makes Language Models Radioactive Tom Sander

Neural Information Processing Systems

Current methods like membership inference or active IP protection either work only in settings where the suspected text is known or do not provide reliable statistical guarantees. We discover that, on the contrary, it is possible to reliably determine if a language model was trained on synthetic data if that data is output by a watermarked LLM.






RankFeat: Rank-1 Feature Removal for Out-of-distribution Detection-Supplementary Material-A Experimental Setup Implementation Details

Neural Information Processing Systems

Table 1 present the evaluation results. We also evaluate our method on the CIFAR benchmark with various model architectures. The best three results are highlighted with red, blue, and cyan. The best two results are highlighted with red and blue . Table 2 compares the performance against all the post hoc baselines.