Goto

Collaborating Authors

 Media


South Park creators' deepfake video startup Deep Voodoo conjures $20M in new funding • TechCrunch

#artificialintelligence

Trey Parker and Matt Stone, creators of South Park and various other media over the years, have raised $20 million to continue work on their professional deepfake studio for creators, Deep Voodoo. The company got its start during the media shutdown of 2020, when the pandemic prevented most travel and on-set productions. Parker and Stone had already begun assembling a AI artist team for a film they were developing, and when COVID intervened they focused on creating the tools for use later. "We stumbled upon this amazing technology and ended up recruiting the best deepfake artists in the world," Stone said in an announcement on Deep Voodoo's site. I've reached out for more info and will update this post if I hear back. The Parker/Stone cachet showed when the company made its public debut alongside no lesser a personage than Kendrick Lamar.


After 35 years of Final Fantasy, what's next for composer Nobuo Uematsu?

Washington Post - Technology News

Uematsu has always been passionate about performing the music he's written for video games onstage. While video game concerts have been taking place in Japan since 1987, when Koichi Sugiyama filled the Suntory Hall in Tokyo with his music from Dragon Quest on the NES, it wasn't until 2003 that Uematsu's music was performed onstage in the West. The success of Thomas Böcker's Symphonic Games Music Concert in Leipzig, Germany, spawned a symphony concert series that awakened Uematsu to the global popularity of Final Fantasy music concerts.


What's next for the metaverse in 2023? - Verdict

#artificialintelligence

The metaverse has been making waves as the next big thing in digital media for months. However, its potential is arguably on every tech-head's mind as we usher in 2023. Research firm GlobalData defines the metaverse as a virtual world where users can share experiences and interact in real-time within simulated scenarios. Its core technologies are vast, but mainly include virtual reality (VR), artificial intelligence (AI) and augmented reality (AR). "Although the metaverse is in the early stages of development, it has the potential to be the next mega-theme in digital media," GlobalData analysts write in a report, "the metaverse could transform how people work, shop, interact, and consume content."


Lensa's viral AI art creations were bound to hypersexualize users - Polygon

#artificialintelligence

This year, it feels like artificial intelligence-generated art has been everywhere. In the summer, many of us entered goofy prompts into DALL-E Mini (now called Craiyon), yielding a series of nine comedically janky AI-generated images. But more recently, there's been a boom of AI-powered apps that can create cool avatars. MyHeritage AI Time Machine generates images of users in historical styles and settings, and AI TikTok filters have become popular for creating anime versions of people. This past week, "magic avatars" from Lensa AI flooded social media platforms like Twitter with illustrative and painterly renderings of people's headshots, as if truly made by magic. These avatars, created using Stable Diffusion -- which allows the AI to "learn" someone's features based off of submitted images -- also opened an ethical can of worms about AI's application.


Audio Denoising for Robust Audio Fingerprinting

arXiv.org Artificial Intelligence

Music discovery services let users identify songs from short mobile recordings. These solutions are often based on Audio Fingerprinting, and rely more specifically on the extraction of spectral peaks in order to be robust to a number of distortions. Few works have been done to study the robustness of these algorithms to background noise captured in real environments. In particular, AFP systems still struggle when the signal to noise ratio is low, i.e when the background noise is strong. In this project, we tackle this problematic with Deep Learning. We test a new hybrid strategy which consists of inserting a denoising DL model in front of a peak-based AFP algorithm. We simulate noisy music recordings using a realistic data augmentation pipeline, and train a DL model to denoise them. The denoising model limits the impact of background noise on the AFP system's extracted peaks, improving its robustness to noise. We further propose a novel loss function to adapt the DL model to the considered AFP system, increasing its precision in terms of retrieved spectral peaks. To the best of our knowledge, this hybrid strategy has not been tested before.


Generating music with sentiment using Transformer-GANs

arXiv.org Artificial Intelligence

The field of Automatic Music Generation has seen significant progress thanks to the advent of Deep Learning. However, most of these results have been produced by unconditional models, which lack the ability to interact with their users, not allowing them to guide the generative process in meaningful and practical ways. Moreover, synthesizing music that remains coherent across longer timescales while still capturing the local aspects that make it sound ``realistic'' or ``human-like'' is still challenging. This is due to the large computational requirements needed to work with long sequences of data, and also to limitations imposed by the training schemes that are often employed. In this paper, we propose a generative model of symbolic music conditioned by data retrieved from human sentiment. The model is a Transformer-GAN trained with labels that correspond to different configurations of the valence and arousal dimensions that quantitatively represent human affective states. We try to tackle both of the problems above by employing an efficient linear version of Attention and using a Discriminator both as a tool to improve the overall quality of the generated music and its ability to follow the conditioning signals.


What do LLMs Know about Financial Markets? A Case Study on Reddit Market Sentiment Analysis

arXiv.org Artificial Intelligence

Market sentiment analysis on social media content requires knowledge of both financial markets and social media jargon, which makes it a challenging task for human raters. The resulting lack of high-quality labeled data stands in the way of conventional supervised learning methods. Instead, we approach this problem using semi-supervised learning with a large language model (LLM). Our pipeline generates weak financial sentiment labels for Reddit posts with an LLM and then uses that data to train a small model that can be served in production. We find that prompting the LLM to produce Chain-of-Thought summaries and forcing it through several reasoning paths helps generate more stable and accurate labels, while using a regression loss further improves distillation quality. With only a handful of prompts, the final model performs on par with existing supervised models. Though production applications of our model are limited by ethical considerations, the model's competitive performance points to the great potential of using LLMs for tasks that otherwise require skill-intensive annotation.


Multimodal Hate Speech Detection from Bengali Memes and Texts

arXiv.org Artificial Intelligence

Numerous machine learning (ML) and deep learning (DL)-based approaches have been proposed to utilize textual data from social media for anti-social behavior analysis like cyberbullying, fake news detection, and identification of hate speech mainly for highly-resourced languages such as English. However, despite having a lot of diversity and millions of native speakers, some languages like Bengali are under-resourced, which is due to a lack of computational resources for natural language processing (NLP). Similar to other languages, Bengali social media contents also include images along with texts (e.g., multimodal memes are posted by embedding short texts into images on Facebook). Therefore, only the textual data is not enough to judge them since images might give extra context to make a proper judgement. This paper is about hate speech detection from multimodal Bengali memes and texts. We prepared the only multimodal hate speech dataset for-a-kind of problem for Bengali, which we use to train state-of-the-art neural architectures (e.g., Bi-LSTM/Conv-LSTM with word embeddings, ConvNets + pre-trained language models, e.g., monolingual Bangla BERT, multilingual BERT-cased/uncased, and XLM-RoBERTa) to jointly analyze textual and visual information for hate speech detection. Conv-LSTM and XLM-RoBERTa models performed best for texts, yielding F1 scores of 0.78 and 0.82, respectively. As of memes, ResNet-152 and DenseNet-161 models yield F1 scores of 0.78 and 0.79, respectively. As for multimodal fusion, XLM-RoBERTa + DenseNet-161 performed the best, yielding an F1 score of 0.83. Our study suggests that text modality is most useful for hate speech detection, while memes are moderately useful.


Multimodal Emotion Recognition among Couples from Lab Settings to Daily Life using Smartwatches

arXiv.org Artificial Intelligence

Couples generally manage chronic diseases together and the management takes an emotional toll on both patients and their romantic partners. Consequently, recognizing the emotions of each partner in daily life could provide an insight into their emotional well-being in chronic disease management. The emotions of partners are currently inferred in the lab and daily life using self-reports which are not practical for continuous emotion assessment or observer reports which are manual, time-intensive, and costly. Currently, there exists no comprehensive overview of works on emotion recognition among couples. Furthermore, approaches for emotion recognition among couples have (1) focused on English-speaking couples in the U.S., (2) used data collected from the lab, and (3) performed recognition using observer ratings rather than partner's self-reported / subjective emotions. In this body of work contained in this thesis (8 papers - 5 published and 3 currently under review in various journals), we fill the current literature gap on couples' emotion recognition, develop emotion recognition systems using 161 hours of data from a total of 1,051 individuals, and make contributions towards taking couples' emotion recognition from the lab which is the status quo, to daily life. This thesis contributes toward building automated emotion recognition systems that would eventually enable partners to monitor their emotions in daily life and enable the delivery of interventions to improve their emotional well-being.


The Gods Themselves: ChatGPT and the Quest for Artificial General Intelligence

#artificialintelligence

In a remarkable feat, ChatGPT gained 1 million users within just five days of its release. For those that missed the news, ChatGPT is a language model for dialogue that interacts in a conversational way. The rapid adoption of ChatGPT highlights the increasing demand for advanced Natural Language Processing (NLP) technology in our society. Not only does it have the potential to revolutionise the way we communicate with machines, but it may also have a significant socio-economic impact. For example, ChatGPT could be used to automate customer service, allowing companies to improve efficiency.