Goto

Collaborating Authors

 Media




Visual Instruction Tuning

Neural Information Processing Systems

Instruction tuning large language models (LLMs) using machine-generated instruction-following data has been shown to improve zero-shot capabilities on new tasks, but the idea is less explored in the multimodal field.


Fans Call on Taylor Swift to 'Do Better' After Accusations of Using AI for Promo Videos

WIRED

Fans Call on Taylor Swift to'Do Better' After Accusations of Using AI for Promo Videos A scavenger hunt campaign to promote Taylor Swift's new album, resulted in a viral #SwiftiesAgainstAI campaign. Fans attend a screening of at a theater in Los Angeles. These were just some of the alleged clues that fans spotted in promo videos for Taylor Swift's new album,, this weekend. They were, to their eyes, telltale indicators that the videos were purportedly made with generative AI . "The first sign that it was AI was that it didn't look great," claims Marcela Lobo, a graphic designer in Brazil who has been a Swift fan since she was 12. "It was wonky, the shadows didn't match, the windows and the painted piano, it looked like shit, basically."



It's True: The Internet Skews the Reality of Women (and Men) in the Workforce

Mother Jones

Age and gender biases are baked into what we see online, a large new study confirms. Get your news from a source that's not owned and controlled by oligarchs. In the 1970s, when researchers asked children to draw a scientist, 99 percent of them drew a man . As this experiment was repeated over 50 years, the number of women drawn increased, and within the past decade, more than half of girls will draw a woman when asked what a scientist looks like. Today, Google search results tend to agree with these children's drawings.


A Manually Annotated Dataset for Instruction-Guided Image Editing

Neural Information Processing Systems

Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are either zero-shot or trained on an automatically synthesized dataset, which contains a high volume of noise. Thus, they still require lots of manual tuning to produce desirable outcomes in practice.