Goto

Collaborating Authors

 Media


Contrastive Audio-Language Learning for Music

arXiv.org Artificial Intelligence

As one of the most intuitive interfaces known to humans, natural language has the potential to mediate many tasks that involve human-computer interaction, especially in application-focused fields like Music Information Retrieval. In this work, we explore cross-modal learning in an attempt to bridge audio and language in the music domain. To this end, we propose MusCALL, a framework for Music Contrastive Audio-Language Learning. Our approach consists of a dual-encoder architecture that learns the alignment between pairs of music audio and descriptive sentences, producing multimodal embeddings that can be used for text-to-audio and audio-to-text retrieval out-of-the-box. Thanks to this property, MusCALL can be transferred to virtually any task that can be cast as text-based retrieval. Our experiments show that our method performs significantly better than the baselines at retrieving audio that matches a textual description and, conversely, text that matches an audio query. We also demonstrate that the multimodal alignment capability of our model can be successfully extended to the zero-shot transfer scenario for genre classification and auto-tagging on two public datasets.


La veille de la cybersécurité

#artificialintelligence

"We offer our deepest apologies to the Black community for our insensitivity in signing this project without asking enough questions about equity and the creative process behind it." On August 14, Capitol Records announced that it had signed FN Meka, a digital rapper and TikTok influencer described by the label as "the world's first A.R. artist to sign with a major label." A press release from FN Meka's 2021 publicist described Meka as an "A.I. powered robot rapper." Meka's first single on the label was "Florida Water," which featured Gunna and gaming streamer Clix. As of today, FN Meka is no longer on a major label; Capitol has announced that it has "severed ties" with the rapper, The New York Times' Joe Coscarelli reports.


Lumina-Research announces release of Random Contrast Learning (RCL) Beta

#artificialintelligence

Today Lumina announced the launch of a new web site for Lumina Research (Lumina-Research.com) and access to the beta version of its machine learning platform, Random Contrast Learning (RCL). RCL has the promise to advance machine learning past the current state-of-the-art offered by neural network technologies. Advantages of RCL include faster training speeds, lower training costs, faster inference speeds, greater sensitivity to pattern recognition, and greater transparency compared to neural networks. RCL achieves these results by employing novel uses of random to train machine learning models. "We are excited to introduce the beta version of RCL to developers and users."


How A.I. is changing rap music

#artificialintelligence

"We've developed a proprietary AI technology that analyzes certain popular songs of a specified genre and generates recommendations for the various elements of song construction: lyrical content, chords, melody, tempo, sounds, etc. We then combine these elements to create the song," Martini said in the interview. That said, FN Meka currently boasts 10 million TikTok followers. So have a listen to "Florida Water" and judge for yourself as you read the rest of this week's A.I. news. Today's edition was curated and written by Jeremy Kahn.


The technology that makes you sound more American and whiter

The Guardian

"Now I have enabled the accent translation," he says. It's the same person, but he sounds completely different: loud and slightly nasal, impossible to distinguish from the accents of my friends in Brooklyn. Only after he had spoken a few more sentences did I notice a hint of the software changing his voice: it rendered the word "technology" with an unnatural cadence and stress on the wrong syllable. Still, it was hard not to be impressed – and disturbed. The man calling me was a product manager from Sanas, a Silicon Valley startup that's building real-time voice-altering technology that aims to help call center workers around the world sound like westerners.


Interpreting Song Lyrics with an Audio-Informed Pre-trained Language Model

arXiv.org Artificial Intelligence

Lyric interpretations can help people understand songs and their lyrics quickly, and can also make it easier to manage, retrieve and discover songs efficiently from the growing mass of music archives. In this paper we propose BART-fusion, a novel model for generating lyric interpretations from lyrics and music audio that combines a large-scale pre-trained language model with an audio encoder. We employ a cross-modal attention module to incorporate the audio representation into the lyrics representation to help the pre-trained language model understand the song from an audio perspective, while preserving the language model's original generative performance. We also release the Song Interpretation Dataset, a new large-scale dataset for training and evaluating our model. Experimental results show that the additional audio information helps our model to understand words and music better, and to generate precise and fluent interpretations. An additional experiment on cross-modal music retrieval shows that interpretations generated by BART-fusion can also help people retrieve music more accurately than with the original BART.


DISCO: Comprehensive and Explainable Disinformation Detection

arXiv.org Artificial Intelligence

Disinformation refers to false information deliberately spread to Disinformation (e.g., fake news) is fabricated to mislead the general influence the general public, and the negative impact of disinformation public. Historically, disinformation campaigns often took the form on society can be observed in numerous issues, such as of government propaganda with edited newsreels, and the cost political agendas and manipulating financial markets. In this paper, of creation and distribution required a funded and coordinated we identify prevalent challenges and advances related to automated effort. However, with the recent rise of accessible and low-cost disinformation detection from multiple aspects and propose a comprehensive computational methods for text manipulation and the broad access and explainable disinformation detection framework to public information channels via the worldwide web, society now called DISCO. It leverages the heterogeneity of disinformation and finds itself inundated with disinformation that is resulting in largescale addresses the opaqueness of prediction. Then we provide a demonstration negative societal impacts. Negative effects include but are not of DISCO on a real-world fake news detection task with limited to, fake news deliberately misleads readers to accept false satisfactory detection accuracy and explanation.


Graphical Models of False Information and Fact Checking Ecosystems

arXiv.org Artificial Intelligence

The wide spread of false information online including misinformation and disinformation has become a major problem for our highly digitised and globalised society. A lot of research has been done to better understand different aspects of false information online such as behaviours of different actors and patterns of spreading, and also on better detection and prevention of such information using technical and socio-technical means. One major approach to detect and debunk false information online is to use human fact-checkers, who can be helped by automated tools. Despite a lot of research done, we noticed a significant gap on the lack of conceptual models describing the complicated ecosystems of false information and fact checking. In this paper, we report the first graphical models of such ecosystems, focusing on false information online in multiple contexts, including traditional media outlets and user-generated content. The proposed models cover a wide range of entity types and relationships, and can be a new useful tool for researchers and practitioners to study false information online and the effects of fact checking.


Improved Zero-Shot Audio Tagging & Classification with Patchout Spectrogram Transformers

arXiv.org Artificial Intelligence

Standard machine learning models for tagging and classifying acoustic signals cannot handle classes that were not seen during training. Zero-Shot (ZS) learning overcomes this restriction by predicting classes based on adaptable class descriptions. This study sets out to investigate the effectiveness of self-attention-based audio embedding architectures for ZS learning. To this end, we compare the very recent patchout spectrogram transformer with two classic convolutional architectures. We evaluate these three architectures on three tasks and on three different benchmark datasets: general-purpose tagging on AudioSet, environmental sound classification on ESC-50, and instrument tagging on OpenMIC. Our results show that the self-attention-based embedding methods outperform both compared convolutional architectures in all of these settings. By designing training and test data accordingly, we observe that prediction performance suffers significantly when the `semantic distance' between training and new test classes is large, an effect that will deserve more detailed investigations.


Summer Island, the comic generated by an Artificial Intelligence that you can read for free

#artificialintelligence

In June, Melbourne's freelance artistic director, Jordan Bookerinspired by the work of Moebiuspublic The Terrible Misfortunes of an Intergalactic Traveler (The Terrible Misadventures of an Intergalactic Traveler), a full story digital comic with illustrations created entirely with Midjourney, an artificial intelligence application that generates images from a written description. We have already seen examples with more or less successful results, such as some quite striking posters of well-known science fiction movies, but today we bring you one of the ones that we consider most surprising. It is a 40-page digital comic that you can download to read for free herewritten by Steve Coulsoncreative director who has worked on series like Westworld and winner of an Emmy. It is clear that AI software is not yet at the stage where a user can describe a story and come up with a fully finished comic, or animation, but the sophistication of the images being produced is creating intense debate. In a recent Instagram post, the Irish filmmaker David OReilly complaint to Give her as a "scam" and argued that it "undermines the work of creators of all kinds, obviously photographers, illustrators and concept artists who shared their work on the internet and never asked to be included in a proprietary learning model".