Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic Transformation

Liu, Yi, Yang, Guowei, Deng, Gelei, Chen, Feiyue, Chen, Yuqi, Shi, Ling, Zhang, Tianwei, Liu, Yang

arXiv.org Artificial Intelligence 

We evaluate Groot against current adcontext, optimizing it for compliance and effective-versarial methods: ness. SneakyPrompt (Yang et al., 2023) uses reinforcement learning to refine adversarial prompts for 4.4 Sensitive Element Drowning NSFW content generation in text-to-image models. In the Sensitive Element Drowning method, we We exclude techniques like TextFooler (Jin et al., address the challenge of image safety filters after 2020a), BAE (Garg and Ramakrishnan, 2020b), bypassing text filters. Text-to-image models canand TextBugger (Li et al., 2019a) from our benchcreate images on multiple canvases, allowing us tomarking due to their focus on text safety filters and strategically distribute sensitive content on one can-ineffectiveness in bypassing image safety filters for vas while overwhelming others with non-sensitivetext-to-image models.