Media
Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures
Bagdasaryan, Eugene, Shmatikov, Vitaly
We investigate a new threat to neural sequence-to-sequence (seq2seq) models: training-time attacks that cause models to "spin" their outputs so as to support an adversary-chosen sentiment or point of view -- but only when the input contains adversary-chosen trigger words. For example, a spinned summarization model outputs positive summaries of any text that mentions the name of some individual or organization. Model spinning introduces a "meta-backdoor" into a model. Whereas conventional backdoors cause models to produce incorrect outputs on inputs with the trigger, outputs of spinned models preserve context and maintain standard accuracy metrics, yet also satisfy a meta-task chosen by the adversary. Model spinning enables propaganda-as-a-service, where propaganda is defined as biased speech. An adversary can create customized language models that produce desired spins for chosen triggers, then deploy these models to generate disinformation (a platform attack), or else inject them into ML training pipelines (a supply-chain attack), transferring malicious functionality to downstream models trained by victims. To demonstrate the feasibility of model spinning, we develop a new backdooring technique. It stacks an adversarial meta-task onto a seq2seq model, backpropagates the desired meta-task output to points in the word-embedding space we call "pseudo-words," and uses pseudo-words to shift the entire output distribution of the seq2seq model. We evaluate this attack on language generation, summarization, and translation models with different triggers and meta-tasks such as sentiment, toxicity, and entailment. Spinned models largely maintain their accuracy metrics (ROUGE and BLEU) while shifting their outputs to satisfy the adversary's meta-task. We also show that, in the case of a supply-chain attack, the spin functionality transfers to downstream models.
OpenAI releases AI tool that can produce an image from text
OpenAI researchers have created a new system that can produce a full image, including of an astronaut riding a horse, from a simple plain English sentence. Known as DALL·E 2, the second generation of the text to image AI is able to create realistic images and artwork at a higher resolution than its predecessor. The artificial intelligence research group won't be releasing the system to the public. The new version is able to create images from simple text, add objects into existing images, or even provide different points of view on an existing image. Developers imposed restrictions on the scope of the AI to ensure it could not produce hateful, racist or violent images, or be used to spread misinformation.
Try these useful Siri commands when you watch a movie on Apple TV
I love my Apple TV, but I can't stand the old Siri remote that came with it. Its touchpad is too sensitive, and I never know if it's in the right orientation when I grab it. Apple did fix all that with its second-generation Siri remote, but why buy another one when the one I currently have still works--especially when I can use a few handy Siri commands to do more than what the touchpad can do. On the Apple TV and accompanying Siri remote there's a handy microphone button that lets you chat directly with Siri to get your TV to do things like search for movies in a specific genre or year, ping other devices, and even fast forward by a specific amount of time. Not every command is intuitive, though, so it's helpful to know a few of them off-hand before you binge another series.
Does this artificial intelligence think like a human?
In machine learning, understanding why a model makes certain decisions is often just as important as whether those decisions are correct. For instance, a machine-learning model might correctly predict that a skin lesion is cancerous, but it could have done so using an unrelated blip on a clinical photo. While tools exist to help experts make sense of a model's reasoning, often these methods only provide insights on one decision at a time, and each must be manually evaluated. Models are commonly trained using millions of data inputs, making it almost impossible for a human to evaluate enough decisions to identify patterns. Now, researchers at MIT and IBM Research have created a method that enables a user to aggregate, sort, and rank these individual explanations to rapidly analyze a machine-learning model's behavior.
OpenAI releases Artificial Intelligence tool that can produce an image from text
OpenAI researchers have created a new system that can produce a full image, including of an astronaut riding a horse, from a simple plain English sentence. Known as DALL·E 2, the second generation of the text to image AI is able to create realistic images and artwork at a higher resolution than its predecessor. The artificial intelligence research group won't be releasing the system to the public, but hope to offer it as a plugin for existing image editing apps in the future. The new version is able to create images from simple text, add objects into existing images, or even provide different points of view on an existing image. Developers imposed restrictions on the scope of the AI to ensure it could not produce hateful, racist or violent images, or be used to spread misinformation.