Media
Microsoft's AI tool can turn photos into realistic videos of people talking and singing
Microsoft Research Asia has unveiled a new experimental AI tool called VASA-1 that can take a still image of a person -- or the drawing of one -- and an existing audio file to create a lifelike talking face out of them in real time. It has the ability to generate facial expressions and head motions for an existing still image and the appropriate lip movements to match a speech or a song. The researchers uploaded a ton of examples on the project page, and the results look good enough that they could fool people into thinking that they're real. While the lip and head motions in the examples could still look a bit robotic and out of sync upon closer inspection, it's still clear that the technology could be misused to easily and quickly create deepfake videos of real people. The researchers themselves are aware of that potential and have decided not to release "an online demo, API, product, additional implementation details, or any related offerings" until they're sure that their technology "will be used responsibly and in accordance with proper regulations."
Do "English" Named Entity Recognizers Work Well on Global Englishes?
Shan, Alexander, Bauer, John, Carlson, Riley, Manning, Christopher
The vast majority of the popular English named entity recognition (NER) datasets contain American or British English data, despite the existence of many global varieties of English. As such, it is unclear whether they generalize for analyzing use of English globally. To test this, we build a newswire dataset, the Worldwide English NER Dataset, to analyze NER model performance on low-resource English variants from around the world. We test widely used NER toolkits and transformer models, including models using the pre-trained contextual models RoBERTa and ELECTRA, on three datasets: a commonly used British English newswire dataset, CoNLL 2003, a more American focused dataset OntoNotes, and our global dataset. All models trained on the CoNLL or OntoNotes datasets experienced significant performance drops-over 10 F1 in some cases-when tested on the Worldwide English dataset. Upon examination of region-specific errors, we observe the greatest performance drops for Oceania and Africa, while Asia and the Middle East had comparatively strong performance. Lastly, we find that a combined model trained on the Worldwide dataset and either CoNLL or OntoNotes lost only 1-2 F1 on both test sets.
Movie101v2: Improved Movie Narration Benchmark
Yue, Zihao, Zhang, Yepeng, Wang, Ziheng, Jin, Qin
Automatic movie narration targets at creating video-aligned plot descriptions to assist visually impaired audiences. It differs from standard video captioning in that it requires not only describing key visual details but also inferring the plots developed across multiple movie shots, thus posing unique and ongoing challenges. To advance the development of automatic movie narrating systems, we first revisit the limitations of existing datasets and develop a large-scale, bilingual movie narration dataset, Movie101v2. Second, taking into account the essential difficulties in achieving applicable movie narration, we break the long-term goal into three progressive stages and tentatively focus on the initial stages featuring understanding within individual clips. We also introduce a new narration assessment to align with our staged task goals. Third, using our new dataset, we baseline several leading large vision-language models, including GPT-4V, and conduct in-depth investigations into the challenges current models face for movie narration generation. Our findings reveal that achieving applicable movie narration generation is a fascinating goal that requires thorough research.
Music Consistency Models
Fei, Zhengcong, Fan, Mingyuan, Huang, Junshi
Consistency models have exhibited remarkable capabilities in facilitating efficient image/video generation, enabling synthesis with minimal sampling steps. It has proven to be advantageous in mitigating the computational burdens associated with diffusion models. Nevertheless, the application of consistency models in music generation remains largely unexplored. To address this gap, we present Music Consistency Models (\texttt{MusicCM}), which leverages the concept of consistency models to efficiently synthesize mel-spectrogram for music clips, maintaining high quality while minimizing the number of sampling steps. Building upon existing text-to-music diffusion models, the \texttt{MusicCM} model incorporates consistency distillation and adversarial discriminator training. Moreover, we find it beneficial to generate extended coherent music by incorporating multiple diffusion processes with shared constraints. Experimental results reveal the effectiveness of our model in terms of computational efficiency, fidelity, and naturalness. Notable, \texttt{MusicCM} achieves seamless music synthesis with a mere four sampling steps, e.g., only one second per minute of the music clip, showcasing the potential for real-time application.
A Survey on the Memory Mechanism of Large Language Model based Agents
Zhang, Zeyu, Bo, Xiaohe, Ma, Chen, Li, Rui, Chen, Xu, Dai, Quanyu, Zhu, Jieming, Dong, Zhenhua, Wen, Ji-Rong
Large language model (LLM) based agents have recently attracted much attention from the research and industry communities. Compared with original LLMs, LLM-based agents are featured in their self-evolving capability, which is the basis for solving real-world problems that need long-term and complex agent-environment interactions. The key component to support agent-environment interactions is the memory of the agents. While previous studies have proposed many promising memory mechanisms, they are scattered in different papers, and there lacks a systematical review to summarize and compare these works from a holistic perspective, failing to abstract common and effective designing patterns for inspiring future studies. To bridge this gap, in this paper, we propose a comprehensive survey on the memory mechanism of LLM-based agents. In specific, we first discuss ''what is'' and ''why do we need'' the memory in LLM-based agents. Then, we systematically review previous studies on how to design and evaluate the memory module. In addition, we also present many agent applications, where the memory module plays an important role. At last, we analyze the limitations of existing work and show important future directions. To keep up with the latest advances in this field, we create a repository at \url{https://github.com/nuster1128/LLM_Agent_Memory_Survey}.
The Taylor Swift Album Leak's Big AI Problem
On Thursday, Taylor Swift did a very Taylor Swift thing: She posted an Instagram story with a link to buy "Fortnight," the first single off of her new album, The Tortured Poets Department. It was cute, maybe even unnecessary. Taylor Swift is one of the biggest recording artists in the world. She announced TTPD in February while accepting the Grammy for best pop vocal album for her last record, Midnights. Swift sold 19 million albums in the US alone last year; she doesn't have to post IG stories about a new single.
Israeli missiles hit site in Iran, media report says
Israeli missiles have hit a site in Iran, ABC News reported late on Thursday, citing a U.S. official, days after Iran launched a drone strike on Israel in response to an attack at the Iranian embassy in Syria. Iran's Fars news agency said an explosion was heard at an airport in the Iranian city of Isafahan, but the cause was not immediately known. Several Iranian nuclear sites are located in Isfahan province, including Natanz, the centerpiece of Iran's uranium enrichment program. Several flights were diverted over Iranian airspace, CNN reported. Over the weekend, Iran launched hundreds of drones and missiles in a retaliatory strike after a suspected Israeli strike on its embassy compound in Syria.
Note: Harnessing Tellurium Nanoparticles in the Digital Realm Plasmon Resonance, in the Context of Brewster's Angle and the Drude Model for Fake News Adsorption in Incomplete Information Games
This note explores the innovative application of soliton theory and plasmonic phenomena in modeling user behavior and engagement within digital health platforms. By introducing the concept of soliton solutions, we present a novel approach to understanding stable patterns of health improvement behaviors over time. Additionally, we delve into the role of tellurium nanoparticles and their plasmonic properties in adsorbing fake news, thereby influencing user interactions and engagement levels. Through a theoretical framework that combines nonlinear dynamics with the unique characteristics of tellurium nanoparticles, we aim to provide new insights into the dynamics of user engagement in digital health environments. Our analysis highlights the potential of soliton theory in capturing the complex, nonlinear dynamics of user behavior, while the application of plasmonic phenomena offers a promising avenue for enhancing the sensitivity and effectiveness of digital health platforms. This research ventures into an uncharted territory where optical phenomena such as Brewster's Angle and Snell's Law, along with the concept of spin solitons, are metaphorically applied to address the challenge of fake news dissemination. By exploring the analogy between light refraction, reflection, and the propagation of information in digital platforms, we unveil a novel perspective on how the 'angle' at which information is presented can significantly affect its acceptance and spread. Additionally, we propose the use of tellurium nanoparticles to manage 'information waves' through mechanisms akin to plasmonic resonance and soliton dynamics. This theoretical exploration aims to bridge the gap between physical sciences and digital communication, offering insights into the development of strategies for mitigating misinformation.
Plasmon Resonance Model: Investigation of Analysis of Fake News Diffusion Model with Third Mover Intervention Using Soliton Solution in Non-Complete Information Game under Repeated Dilemma Condition
In this study, we attempt to model the prominent problem of fake news diffusion in modern society using the framework of incomplete information games and nonlinear partial differential equations. In particular, we focus on the plasmon resonance phenomenon, in which fake news diffuses rapidly under certain conditions and causes significant social impact, and aim to theoretically elucidate its mechanism. We also incorporate the concepts of first movers, second movers, and third movers in game theory to explore how their strategies affect the dynamics of fake news diffusion. The proliferation of fake news is a complex process in which truth and misinformation intersect, and its effects reach across political, economic, and social strata. To address this issue, it is essential to understand how fake news is widely accepted and shared. In this study, we liken this diffusion process to the concept of plasmon resonance in physics to model the phenomenon of the rapid amplification of fake Figure 1: Comparison of Third Mover Soliton Solution and news within a particular social group. Plasmon resonance is Parabolic Strategies under Plasmon Influence a resonance phenomenon that occurs when electron density waves interact with light on a metal surface.