Media
A Bruce Willis deepfake will appear in his stead for future film projects
Bruce Willis may have retired from acting following a diagnosis of aphasia, but a version of him will live on in future projects. Last year, the actor's "digital twin" appeared in an ad for a Russian telecom created by a company called Deepcake. Now, it's being reported that he sold his rights for future film, advertising and other projects to Deepcake, according to the company's website and The Telegraph. Engineers created the digital double drawing from content in Die Hard and Fifth Element, when Willis was 32 and 42, respectively. However, Willis's estate has final approval on any projects.
'Star Wars' Actor Joins Ukraine's Fight Against Russia As Army Of Drones Ambassador
"Star Wars" star Mark Hamill has been appointed as the ambassador of UNITED24, a fundraising platform supporting the "Army of Drones" project for the benefit of Ukraine. The 71-year-old veteran actor, famous for his role as Luke Skywalker in the galactic franchise, was introduced as the ambassador via an online call with Ukrainian President Volodymyr Zelenskyy Thursday, according to CNN. The president reportedly "expressed his gratitude" toward Hamill, who had been actively showing his support to his home country, which was invaded by Russia earlier this year. "Mark, you have become the first ambassador to help Ukraine raise funds to support its defenders," Zelenskyy said. "For Ukrainians, this means a lot. As in'Star Wars,' good will triumph over evil, and light will overcome darkness. Hamill responded and said he was "honored" to have such a vital role in this "long and unequal fight" between Russia and Ukraine because the latter needed "continuous additional support." "I know for certain that Ukrainians need drones to protect their land, their freedom, and the values of the entire democratic world.
DALL-E is now available to all. NPR put it to work
"A Cubist painting of a mug with the NPR logo on it, on a table next to a old-timey radio" Image generated by DALL-E/OpenAI hide caption An artificial intelligence tool called DALL-E that's stunned with its ability to render text into realistic images is now available to the public. OpenAI, the Silicon Valley research lab behind the program, announced Wednesday it has dropped the waitlist to use the program. Until now, OpenAI released the tool to a select group of users that included academics, artists and journalists. The iterative rollout was designed to curb the potential for bad actors to leverage the tool for disinformation and other harmful uses. The excitement over the invite-only tool had meanwhile inspired an imitation known as DALL-E mini, a limited model in comparison that's not affiliated with OpenAI.
Meta announces AI-based tool for generating video from text
Facebook parent company Meta Thursday announced Make a Video, an online tool that generates short movie clips based on a text description. Why it matters: Text-to-still-image AI systems, including DALL-E 2, Stable Diffusion, and rival projects from Google, Meta and others, have advanced with great speed over the past year. Now, Meta appears to have won the race to extend AI-based content creation to videos. Details: As with the still-image generators, all people have to do is describe something they want to see in a verbal prompt and the system returns a visual result -- in this case, a video. The big picture: AI image generators have wowed observers, but they also raise significant concerns regarding ownership, privacy, the data sets they're based on and the impact their advent may have on artists' income and employment.
AI image generators will help artists, not replace them
For years, artist Steve Coulson wanted to make his own comic. "The problem has always been โ I can't draw," he says. But in 2022, Coulson published a beautiful comic called Summer Island. The 40-page folk-horror story about a sea god festival features detailed illustrations with a coherent visual style-- all created with the help of artificial intelligence. As AI image generators, such as OpenAI's popular DALL-E and DALL-E2, become more widespread, some forecast the death of human artforms.
PART: Pre-trained Authorship Representation Transformer
Huertas-Tato, Javier, Huertas-Garcia, Alvaro, Martin, Alejandro, Camacho, David
Authors writing documents imprint identifying information within their texts: vocabulary, registry, punctuation, misspellings, or even emoji usage. Finding these details is very relevant to profile authors, relating back to their gender, occupation, age, and so on. But most importantly, repeating writing patterns can help attributing authorship to a text. Previous works use hand-crafted features or classification tasks to train their authorship models, leading to poor performance on out-of-domain authors. A better approach to this task is to learn stylometric representations, but this by itself is an open research challenge. In this paper, we propose PART: a contrastively trained model fit to learn \textbf{authorship embeddings} instead of semantics. By comparing pairs of documents written by the same author, we are able to determine the proprietary of a text by evaluating the cosine similarity of the evaluated documents, a zero-shot generalization to authorship identification. To this end, a pre-trained Transformer with an LSTM head is trained with the contrastive training method. We train our model on a diverse set of authors, from literature, anonymous blog posters and corporate emails; a heterogeneous set with distinct and identifiable writing styles. The model is evaluated on these datasets, achieving zero-shot 72.39\% and 86.73\% accuracy and top-5 accuracy respectively on the joint evaluation dataset when determining authorship from a set of 250 different authors. We qualitatively assess the representations with different data visualizations on the available datasets, profiling features such as book types, gender, age, or occupation of the author.
Adversarial Robustness of Representation Learning for Knowledge Graphs
Knowledge graphs represent factual knowledge about the world as relationships between concepts and are critical for intelligent decision making in enterprise applications. New knowledge is inferred from the existing facts in the knowledge graphs by encoding the concepts and relations into low-dimensional feature vector representations. The most effective representations for this task, called Knowledge Graph Embeddings (KGE), are learned through neural network architectures. Due to their impressive predictive performance, they are increasingly used in high-impact domains like healthcare, finance and education. However, are the black-box KGE models adversarially robust for use in domains with high stakes? This thesis argues that state-of-the-art KGE models are vulnerable to data poisoning attacks, that is, their predictive performance can be degraded by systematically crafted perturbations to the training knowledge graph. To support this argument, two novel data poisoning attacks are proposed that craft input deletions or additions at training time to subvert the learned model's performance at inference time. These adversarial attacks target the task of predicting the missing facts in knowledge graphs using KGE models, and the evaluation shows that the simpler attacks are competitive with or outperform the computationally expensive ones. The thesis contributions not only highlight and provide an opportunity to fix the security vulnerabilities of KGE models, but also help to understand the black-box predictive behaviour of KGE models.
K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment
Lee, Jean, Lim, Taejun, Lee, Heejun, Jo, Bogeun, Kim, Yangsok, Yoon, Heegeun, Han, Soyeon Caren
Online hate speech detection has become an important issue due to the growth of online content, but resources in languages other than English are extremely limited. We introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns. The dataset consists of 109k utterances from news comments and provides a multi-label classification using 1 to 4 labels, and handles subjectivity and intersectionality. We evaluate strong baseline experiments on K-MHaS using Korean-BERT-based language models with six different metrics. KR-BERT with a sub-character tokenizer outperforms others, recognizing decomposed characters in each hate speech class.
The Minority Matters: A Diversity-Promoting Collaborative Metric Learning Algorithm
Bao, Shilong, Xu, Qianqian, Yang, Zhiyong, He, Yuan, Cao, Xiaochun, Huang, Qingming
Collaborative Metric Learning (CML) has recently emerged as a popular method in recommendation systems (RS), closing the gap between metric learning and Collaborative Filtering. Following the convention of RS, existing methods exploit unique user representation in their model design. This paper focuses on a challenging scenario where a user has multiple categories of interests. Under this setting, we argue that the unique user representation might induce preference bias, especially when the item category distribution is imbalanced. To address this issue, we propose a novel method called \textit{Diversity-Promoting Collaborative Metric Learning} (DPCML), with the hope of considering the commonly ignored minority interest of the user. The key idea behind DPCML is to include a multiple set of representations for each user in the system. Based on this embedding paradigm, user preference toward an item is aggregated from different embeddings by taking the minimum item-user distance among the user embedding set. Furthermore, we observe that the diversity of the embeddings for the same user also plays an essential role in the model. To this end, we propose a \textit{diversity control regularization} term to accommodate the multi-vector representation strategy better. Theoretically, we show that DPCML could generalize well to unseen test data by tackling the challenge of the annoying operation that comes from the minimum value. Experiments over a range of benchmark datasets speak to the efficacy of DPCML.
Actor Bruce Willis Becomes First Celebrity to Sell Rights to Deepfake Firm
Action movie legend Bruce Willis has just become the first Hollywood actor to sell his rights to the possibility of a "digital twin" to the US firm Deepcake, according to The Telegraph. With the use of deepfake technology, Willis has offered his likeness to be used onscreen for future projects, following his first experience with the digital media manipulation in a commercial for Russian phone service, MegaFon, last year. Deepfake technology allows for the use of a person's likeness to be superimposed over another individual. Through the use of machine learning and AI, it's possible to create a visual and audio "twin" of someone in videos. Though the ability to recreate someone so nearly-flawlessly does raise a few ethical questions, the technology has already been utilized within the Star Wars universe with Rogue One: A Star Wars Story, as well as The Mandalorian Season 2. In 2021, Willis gave permission to Deepcake in order to appear in a commercial, allowing his face to be "digitally transplanted onto another performer."