It's easy to tamper with watermarks from AI-generated text
AI language models work by predicting the next likely word in a sentence, generating one word at a time on the basis of those predictions. Watermarking algorithms for text divide the language model's vocabulary into words on a "green list" and a "red list," and then make the AI model choose words from the green list. The more words in a sentence that are from the green list, the more likely it is that the text was generated by a computer. Humans tend to write sentences that include a more random mix of words. They were able to reverse-engineer the watermarks by using an API to access the AI model with the watermark applied and prompting it many times, says Staab. The responses allow the attacker to "steal" the watermark by building an approximate model of the watermarking rules.
Mar-29-2024, 14:51:24 GMT
- Country:
- North America > United States
- Maryland (0.06)
- Europe > Switzerland
- North America > United States
- Industry:
- Information Technology > Security & Privacy (0.62)
- Technology: