Media
A Survey on Music Generation from Single-Modal, Cross-Modal, and Multi-Modal Perspectives
Li, Shuyu, Ji, Shulei, Wang, Zihao, Wu, Songruoyao, Yu, Jiaxing, Zhang, Kejun
Multi-modal music generation, using multiple modalities like text, images, and video alongside musical scores and audio as guidance, is an emerging research area with broad applications. This paper reviews this field, categorizing music generation systems from the perspective of modalities. The review covers modality representation, multi-modal data alignment, and their utilization to guide music generation. Current datasets and evaluation methods are also discussed. Key challenges in this area include effective multi-modal integration, large-scale comprehensive datasets, and systematic evaluation methods. Finally, an outlook on future research directions is provided, focusing on creativity, efficiency, multi-modal alignment, and evaluation.
Using generative AI will 'neither help nor harm the chances of achieving' Oscar nominations
The Academy of Motion Picture Arts and Sciences has decide that its official stance towards AI-use in films is to take no stance at all, according to a statement the organization shared outlining changes to voting for the 98th Oscars. The issue of award-nominated films using AI was first raised in 2024 when the productions behind Best Picture nominees The Brutalist and Emilia Pérez admitted to using the tech to alter performances. "With regard to Generative Artificial Intelligence and other digital tools used in the making of the film, the tools neither help nor harm the chances of achieving a nomination, " AMPAS writes. "The Academy and each branch will judge the achievement, taking into account the degree to which a human was at the heart of the creative authorship when choosing which movie to award." While the organization at least reaffirms that human involvement is their primary concern, they also don't seem to believe that using AI -- potentially trained on the ill-gotten work of their membership -- is an existential problem.
Subtitling Your Life
A little over thirty years ago, when he was in his mid-forties, my friend David Howorth lost all hearing in his left ear, a calamity known as single-sided deafness. "It happened literally overnight," he said. "My doctor told me, 'We really don't understand why.' " At the time, he was working as a litigator in the Portland, Oregon, office of a large law firm. His hearing loss had no impact on his job--"In a courtroom, you can get along fine with one ear"--but other parts of his life were upended. The brain pinpoints sound sources in part by analyzing minute differences between left-ear and right-ear arrival times, the same process that helps bats and owls find prey they can't see.
Russia-Ukraine war: List of key events, day 1,152
At least three blasts were heard in the Russian-controlled Donetsk region in eastern Ukraine amid an Easter ceasefire declared by Moscow, Russian state news agency TASS reported, citing local "operative services." Ukraine's forces reported nearly 3,000 violations of Russia's own ceasefire pledge, Ukrainian President Volodymyr Zelenskyy said, adding that Kyiv's forces were instructed to mirror the Russian Army's actions. Russia's Ministry of Defence said Ukraine had broken the Easter ceasefire declared by the Kremlin more than a thousand times, claiming that Ukrainian forces shot at Russian positions 444 times. The ministry also said Kremlin forces encountered more than 900 Ukrainian drone attacks during this time. At least three blasts were heard in the Russian-controlled Donetsk region in eastern Ukraine amid an Easter ceasefire declared by Moscow, Russian state news agency TASS reported, citing local "operative services."
The Last of Us season two 'Through the Valley' recap: Well, that happened
HBO's The Last of Us showed viewers in season one that it would lean heavily on the source video games for major plot points and general direction of the season while expanding on the universe, and season two has followed that to the most extreme end possible. Episode two sees Tommy and Maria lead the town of Jackson Hole against a massive wave of Infected, the likes of which we haven't seen in the show (or video games) yet. This was a complete invention for the show, one that gives the episode Game of Thrones vibes, or calls to mind a battle like the siege of Helm's Deep in Lord of the Rings: The Two Towers. It's epic in scale, with the overmatched defenders showing their skill and bravery against overwhelming odds; there is loss and pain but the good guys eventually triumph. That mass-scale battle is paired with the most intimate and brutal violence we've seen in the entire series so far, as Joel's actions finally catch up with him.
Learning to Attribute with Attention
Cohen-Wang, Benjamin, Chuang, Yung-Sung, Madry, Aleksander
Given a sequence of tokens generated by a language model, we may want to identify the preceding tokens that influence the model to generate this sequence. Performing such token attribution is expensive; a common approach is to ablate preceding tokens and directly measure their effects. To reduce the cost of token attribution, we revisit attention weights as a heuristic for how a language model uses previous tokens. Naive approaches to attribute model behavior with attention (e.g., averaging attention weights across attention heads to estimate a token's influence) have been found to be unreliable. To attain faithful attributions, we propose treating the attention weights of different attention heads as features. This way, we can learn how to effectively leverage attention weights for attribution (using signal from ablations). Our resulting method, Attribution with Attention (AT2), reliably performs on par with approaches that involve many ablations, while being significantly more efficient. To showcase the utility of AT2, we use it to prune less important parts of a provided context in a question answering setting, improving answer quality. We provide code for AT2 at https://github.com/MadryLab/AT2 .
Exploring Multimodal Prompt for Visualization Authoring with Large Language Models
Wen, Zhen, Weng, Luoxuan, Tang, Yinghao, Zhang, Runjin, Liu, Yuxin, Pan, Bo, Zhu, Minfeng, Chen, Wei
Recent advances in large language models (LLMs) have shown great potential in automating the process of visualization authoring through simple natural language utterances. However, instructing LLMs using natural language is limited in precision and expressiveness for conveying visualization intent, leading to misinterpretation and time-consuming iterations. To address these limitations, we conduct an empirical study to understand how LLMs interpret ambiguous or incomplete text prompts in the context of visualization authoring, and the conditions making LLMs misinterpret user intent. Informed by the findings, we introduce visual prompts as a complementary input modality to text prompts, which help clarify user intent and improve LLMs' interpretation abilities. To explore the potential of multimodal prompting in visualization authoring, we design VisPilot, which enables users to easily create visualizations using multimodal prompts, including text, sketches, and direct manipulations on existing visualizations. Through two case studies and a controlled user study, we demonstrate that VisPilot provides a more intuitive way to create visualizations without affecting the overall task efficiency compared to text-only prompting approaches. Furthermore, we analyze the impact of text and visual prompts in different visualization tasks. Our findings highlight the importance of multimodal prompting in improving the usability of LLMs for visualization authoring. We discuss design implications for future visualization systems and provide insights into how multimodal prompts can enhance human-AI collaboration in creative visualization tasks. All materials are available at https://OSF.IO/2QRAK.
Doctor Who 'Lux' review: Hope can change the world
It's an interesting time to be a long-running science fantasy media property in the streaming TV age. Star Trek is in the grip of an existential crisis as it (wrongly) fears it's too old-aged to be relevant. Star Wars became a battlefield in the culture war and, to duck all future bad faith criticism, gave us The Rise of Skywalker. And then there's Doctor Who, which is somehow managing to plough a 62-year furrow and still fill it with original ideas. This week the Doctor and Belinda go up against a sentient cartoon holding the patrons of a 1950s cinema hostage.
Filmmaker James Cameron on penguins, arctic cold, and lowlight cameras
James Cameron wasn't near the penguins this time around, but he is extremely familiar with their environment. "When I went to Antarctica myself, I had a Nikon still camera adapted to the cold with special lubricants," he tells Popular Science. "I went to the South Pole and the film shattered in my hand when I tried to change it. I took a video camera, I wrapped it in a heating pack and it [died] in two minutes. I have a good sense of what it takes to take conventional equipment into that environment and survive."
FoxNews AI Newsletter: 'Terminator' director James Cameron flip-flops on AI, says Hollywood is 'looking at it
Reachy 2 is touted as a "lab partner for the AI era." Director James Cameron attends the "Avatar: The Way Of Water" World Premiere at Odeon Luxe Leicester Square in 2022 in London, England. 'I'LL BE BACK': James Cameron's stance on artificial intelligence has evolved over the past few years, and he feels Hollywood needs to embrace it in a few different ways. MADE IN AMERICA: Nvidia on Monday announced plans to manufacture its artificial intelligence supercomputers entirely in the U.S. for the first time. RIDEABLE 4-LEGGED ROOT: Kawasaki Heavy Industries has introduced something that feels straight out of a video game: CORLEO, a hydrogen-powered, four-legged robot prototype designed to be ridden by humans.