Media
Artificial intelligence means anyone can cast Hollywood stars in their own films
For years, the only way to create a blockbuster film featuring a Hollywood star and dazzling special effects was at a major studio. The Hollywood giants were the ones that could afford to pay celebrities millions of dollars and license sophisticated software to produce elaborate, special effects-laden films. That's all about to change, and the public is getting a preview thanks to artificial intelligence (AI) tools like OpenAI's DALL-E and Midjourney. Both tools use images scraped from the internet and select datasets like LAION to train their AI models to reconstruct similar yet wholly original imagery using text prompts. The AI images, which vary from photographic realism to mimicking the styles of famous artists, can be generated in as little as 20 to 30 seconds, often producing results that would take a human hours to produce.
Photographer Accurately Recreates his Work with AI Image Generator
A photographer uploaded his fantastical photos to the AI image generator Midjourney and asked it to recreate them with astonishing results. Artificially intelligent (AI) image generators synthesize pictures from text prompts, but photographers can upload their own photos as a guide. After uploading his original pictures to Midjourney he then types in a text prompt that roughly describes his photo. "I've been playing with AI for almost a year now and I'm starting to get the hang of it even though the technology is moving at the speed of light," he tells PetaPixel. "I was curious and decided to try to use AI to replicate my own pictures" Karppinen is a professional commercial photographer but for his own projects shoots all types of "storytelling" images.
Why authorized deepfakes are becoming big for business
Join us on November 9 to learn how to successfully innovate and achieve efficiency by upskilling and scaling citizen developers at the Low-Code/No-Code Summit. "Deepfake implies unauthorized use of synthetic media and generative artificial intelligence -- we are authorized from the get-go," she told VentureBeat. She described the Tel Aviv- and New York-based Hour One as an AI company that has also "built a legal and ethical framework for how to engage with real people to generate their likeness in digital form." It's an important delineation in an era when deepfakes, or synthetic media in which a person in an existing image or video is replaced with someone else's likeness, has gotten a boatload of bad press -- not surprisingly, given deepfakes' longstanding connection to revenge porn and fake news. The term "deepfake" can be traced to a Reddit user in 2017 named "deepfakes" who, along with others in the community, shared videos, many involving celebrity faces swapped onto the bodies of actresses in pornographic videos.
Casual Conversations v2: Designing a large consent-driven dataset to measure algorithmic bias and robustness
Hazirbas, Caner, Bang, Yejin, Yu, Tiezheng, Assar, Parisa, Porgali, Bilal, Albiero, Vítor, Hermanek, Stefan, Pan, Jacqueline, McReynolds, Emily, Bogen, Miranda, Fung, Pascale, Ferrer, Cristian Canton
Several recent studies [8, 41, 55, 67, 75] propose various learning strategies for AI models to be well-calibrated across all protected subgroups, while others focus on collecting responsible datasets [57, 82, 124] to make sure evaluations of AI models are accurate and algorithmic bias can be measured while promoting data privacy. There has been much criticism regarding the design choice of the publicly used datasets, such as for ImageNet [36, 38, 56, 70]. Discussions are mostly focused on concerns around collecting sensitive data about people without their consent. Casual Conversations v1 [57] was one of the first benchmarks that was designed with permission from participants. However, that dataset has several limitations: samples were collected only in the US, the gender label is limited to three options, and only age and gender labels are self-provided with the permission of the participants.
Vis2Mus: Exploring Multimodal Representation Mapping for Controllable Music Generation
Zhang, Runbang, Zhang, Yixiao, Shao, Kai, Shan, Ying, Xia, Gus
In this study, we explore the representation mapping from the domain of visual arts to the domain of music, with which we can use visual arts as an effective handle to control music generation. Unlike most studies in multimodal representation learning that are purely data-driven, we adopt an analysis-by-synthesis approach that combines deep music representation learning with user studies. Such an approach enables us to discover \textit{interpretable} representation mapping without a huge amount of paired data. In particular, we discover that visual-to-music mapping has a nice property similar to equivariant. In other words, we can use various image transformations, say, changing brightness, changing contrast, style transfer, to control the corresponding transformations in the music domain. In addition, we released the Vis2Mus system as a controllable interface for symbolic music generation.
Optimal Condition Training for Target Source Separation
Tzinis, Efthymios, Wichern, Gordon, Smaragdis, Paris, Roux, Jonathan Le
Recent research has shown remarkable performance in leveraging multiple extraneous conditional and non-mutually exclusive semantic concepts for sound source separation, allowing the flexibility to extract a given target source based on multiple different queries. In this work, we propose a new optimal condition training (OCT) method for single-channel target source separation, based on greedy parameter updates using the highest performing condition among equivalent conditions associated with a given target source. Our experiments show that the complementary information carried by the diverse semantic concepts significantly helps to disentangle and isolate sources of interest much more efficiently compared to single-conditioned models. Moreover, we propose a variation of OCT with condition refinement, in which an initial conditional vector is adapted to the given mixture and transformed to a more amenable representation for target source extraction. We showcase the effectiveness of OCT on diverse source separation experiments where it improves upon permutation invariant models with oracle assignment and obtains state-of-the-art performance in the more challenging task of text-based source separation, outperforming even dedicated text-only conditioned models.
DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question Answering
Neeman, Ella, Aharoni, Roee, Honovich, Or, Choshen, Leshem, Szpektor, Idan, Abend, Omri
Question answering models commonly have access to two sources of "knowledge" during inference time: (1) parametric knowledge - the factual knowledge encoded in the model weights, and (2) contextual knowledge - external knowledge (e.g., a Wikipedia passage) given to the model to generate a grounded answer. Having these two sources of knowledge entangled together is a core issue for generative QA models as it is unclear whether the answer stems from the given non-parametric knowledge or not. This unclarity has implications on issues of trust, interpretability and factuality. In this work, we propose a new paradigm in which QA models are trained to disentangle the two sources of knowledge. Using counterfactual data augmentation, we introduce a model that predicts two answers for a given question: one based on given contextual knowledge and one based on parametric knowledge. Our experiments on the Natural Questions dataset show that this approach improves the performance of QA models by making them more robust to knowledge conflicts between the two knowledge sources, while generating useful disentangled answers.
Learning Sparse Analytic Filters for Piano Transcription
Cwitkowitz, Frank, Heydari, Mojtaba, Duan, Zhiyao
In recent years, filterbank learning has become an increasingly popular strategy for various audio-related machine learning tasks. This is partly due to its ability to discover task-specific audio characteristics which can be leveraged in downstream processing. It is also a natural extension of the nearly ubiquitous deep learning methods employed to tackle a diverse array of audio applications. In this work, several variations of a frontend filterbank learning module are investigated for piano transcription, a challenging low-level music information retrieval task. We build upon a standard piano transcription model, modifying only the feature extraction stage. The filterbank module is designed such that its complex filters are unconstrained 1D convolutional kernels with long receptive fields. Additional variations employ the Hilbert transform to render the filters intrinsically analytic and apply variational dropout to promote filterbank sparsity. Transcription results are compared across all experiments, and we offer visualization and analysis of the filterbanks.
GREENER: Graph Neural Networks for News Media Profiling
Panayotov, Panayot, Shukla, Utsav, Sencar, Husrev Taha, Nabeel, Mohamed, Nakov, Preslav
We study the problem of profiling news media on the Web with respect to their factuality of reporting and bias. This is an important but under-studied problem related to disinformation and "fake news" detection, but it addresses the issue at a coarser granularity compared to looking at an individual article or an individual claim. This is useful as it allows to profile entire media outlets in advance. Unlike previous work, which has focused primarily on text (e.g.,~on the text of the articles published by the target website, or on the textual description in their social media profiles or in Wikipedia), here our main focus is on modeling the similarity between media outlets based on the overlap of their audience. This is motivated by homophily considerations, i.e.,~the tendency of people to have connections to people with similar interests, which we extend to media, hypothesizing that similar types of media would be read by similar kinds of users. In particular, we propose GREENER (GRaph nEural nEtwork for News mEdia pRofiling), a model that builds a graph of inter-media connections based on their audience overlap, and then uses graph neural networks to represent each medium. We find that such representations are quite useful for predicting the factuality and the bias of news media outlets, yielding improvements over state-of-the-art results reported on two datasets. When augmented with conventionally used representations obtained from news articles, Twitter, YouTube, Facebook, and Wikipedia, prediction accuracy is found to improve by 2.5-27 macro-F1 points for the two tasks.