Optimizing watermarks for large language models

Wouters, Bram

arXiv.org Artificial Intelligence 

Before a word is generated, the complete vocabulary of the LLM is split in two disjunct lists With the rise of large language models (LLMs) labelled green and red. This split is pseudo-random, where and concerns about potential misuse, watermarks the seed is determined by the previous word(s). Green-list for generative LLMs have recently attracted much words are then sampled with a higher probability than the attention. An important aspect of such watermarks original LLM prescribes, and red-list words with a lower is the trade-off between their identifiability probability. A detector with knowledge of the pseudorandom and their impact on the quality of the generated green-red split can count the number of green-list text. This paper introduces a systematic approach words in a text. If this number is larger than one would expect to this trade-off in terms of a multi-objective optimization from a text generated without knowledge of the greenred problem. For a large class of robust, efficient split (e.g., a human-generated text), the null hypothesis watermarks, the associated Pareto optimal is rejected and the text is attributed to an LLM.