Goto

Collaborating Authors

 Media


BERT-like Pre-training for Symbolic Piano Music Classification Tasks

arXiv.org Artificial Intelligence

This article presents a benchmark study of symbolic piano music classification using the masked language modelling approach of the Bidirectional Encoder Representations from Transformers (BERT). Specifically, we consider two types of MIDI data: MIDI scores, which are musical scores rendered directly into MIDI with no dynamics and precisely aligned with the metrical grid notated by its composer and MIDI performances, which are MIDI encodings of human performances of musical scoresheets. With five public-domain datasets of single-track piano MIDI files, we pre-train two 12-layer Transformer models using the BERT approach, one for MIDI scores and the other for MIDI performances, and fine-tune them for four downstream classification tasks. These include two note-level classification tasks (melody extraction and velocity prediction) and two sequence-level classification tasks (style classification and emotion classification). Our evaluation shows that the BERT approach leads to higher classification accuracy than recurrent neural network (RNN)-based baselines.


Logic Query of Thoughts: Guiding Large Language Models to Answer Complex Logic Queries with Knowledge Graphs

arXiv.org Artificial Intelligence

Despite the superb performance in many tasks, large language models (LLMs) bear the risk of generating hallucination or even wrong answers when confronted with tasks that demand the accuracy of knowledge. The issue becomes even more noticeable when addressing logic queries that require multiple logic reasoning steps. On the other hand, knowledge graph (KG) based question answering methods are capable of accurately identifying the correct answers with the help of knowledge graph, yet its accuracy could quickly deteriorate when the knowledge graph itself is sparse and incomplete. It remains a critical challenge on how to integrate knowledge graph reasoning with LLMs in a mutually beneficial way so as to mitigate both the hallucination problem of LLMs as well as the incompleteness issue of knowledge graphs. In this paper, we propose 'Logic-Query-of-Thoughts' (LGOT) which is the first of its kind to combine LLMs with knowledge graph based logic query reasoning. LGOT seamlessly combines knowledge graph reasoning and LLMs, effectively breaking down complex logic queries into easy to answer subquestions. Through the utilization of both knowledge graph reasoning and LLMs, it successfully derives answers for each subquestion. By aggregating these results and selecting the highest quality candidate answers for each step, LGOT achieves accurate results to complex questions. Our experimental findings demonstrate substantial performance enhancements, with up to 20% improvement over ChatGPT.


Uruguayan singer-songwriter Jorge Drexler is embarking on his first tour of Europe

FOX News

Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. MADRID (AP) -- Capitalizing on an impressive moment in Spanish-language music, Uruguayan singer-songwriter Jorge Drexler is embarking on his first tour of Europe. "Even Don Quixote didn't get as far as urban Spanish-speaking music is getting today in the world. You can go everywhere and you will find music that was written in Spanish," he told The Associated Press in an interview.


TCL's first original movie is an absurd-looking, AI-generated love story

Engadget

Many major tech companies, particularly those that operate in the TV hardware business, have dipped their toes into original content. Although it's had its own free, ad-supported TV (FAST) channels for a while, TCL is late to that party. Not for much longer though, as the company is set to release its first special, a short romance movie, on TCLtv this summer. There's just one slight hitch: TCL is using generative AI to make original content for its platform, and early signs do not bode well. The company has released the first trailer for Next Stop Paris, which it's calling "the first AI-powered love story."


Suno AI can generate power ballads about coffee – and jingles for the Guardian. But will it hurt musicians?

The Guardian

Heralded as the ChatGPT for music, Suno AI is the latest iteration of generative artificial intelligence to flood social feeds, wowing users with its (ahem) lyrical prowess. Plug in the musical style you want, a genre and a prompt for lyrics and Suno can spit out a full song for you in a matter of seconds. The business has been around for two years, formulated by a group of machine learning experts in Cambridge who struck an interest in audio, according to a profile in Rolling Stone last month. From the outset, making silly songs is slightly addictive. The lyrics might seem shallow and soulless, but they're also often hilarious.


Detecting AI-Generated Images via CLIP

arXiv.org Artificial Intelligence

As AI-generated image (AIGI) methods become more powerful and accessible, it has become a critical task to determine if an image is real or AI-generated. Because AIGI lack the signatures of photographs and have their own unique patterns, new models are needed to determine if an image is AI-generated. In this paper, we investigate the ability of the Contrastive Language-Image Pre-training (CLIP) architecture, pre-trained on massive internet-scale data sets, to perform this differentiation. We fine-tune CLIP on real images and AIGI from several generative models, enabling CLIP to determine if an image is AI-generated and, if so, determine what generation method was used to create it. We show that the fine-tuned CLIP architecture is able to differentiate AIGI as well or better than models whose architecture is specifically designed to detect AIGI. Our method will significantly increase access to AIGI-detecting tools and reduce the negative effects of AIGI on society, as our CLIP fine-tuning procedures require no architecture changes from publicly available model repositories and consume significantly less GPU resources than other AIGI detection models.


Toward Informal Language Processing: Knowledge of Slang in Large Language Models

arXiv.org Artificial Intelligence

Recent advancement in large language models (LLMs) has offered a strong potential for natural language systems to process informal language. A representative form of informal language is slang, used commonly in daily conversations and online social media. To date, slang has not been comprehensively evaluated in LLMs due partly to the absence of a carefully designed and publicly accessible benchmark. Using movie subtitles, we construct a dataset that supports evaluation on a diverse set of tasks pertaining to automatic processing of slang. For both evaluation and finetuning, we show the effectiveness of our dataset on two core applications: 1) slang detection, and 2) identification of regional and historical sources of slang from natural sentences. We also show how our dataset can be used to probe the output distributions of LLMs for interpretive insights. We find that while LLMs such as GPT-4 achieve good performance in a zero-shot setting, smaller BERT-like models finetuned on our dataset achieve comparable performance. Furthermore, we show that our dataset enables finetuning of LLMs such as GPT-3.5 that achieve substantially better performance than strong zero-shot baselines. Our work offers a comprehensive evaluation and a high-quality benchmark on English slang based on the OpenSubtitles corpus, serving both as a publicly accessible resource and a platform for applying tools for informal language processing.


Mitigating Receiver Impact on Radio Frequency Fingerprint Identification via Domain Adaptation

arXiv.org Artificial Intelligence

Radio Frequency Fingerprint Identification (RFFI), which exploits non-ideal hardware-induced unique distortion resident in the transmit signals to identify an emitter, is emerging as a means to enhance the security of communication systems. Recently, machine learning has achieved great success in developing state-of-the-art RFFI models. However, few works consider cross-receiver RFFI problems, where the RFFI model is trained and deployed on different receivers. Due to altered receiver characteristics, direct deployment of RFFI model on a new receiver leads to significant performance degradation. To address this issue, we formulate the cross-receiver RFFI as a model adaptation problem, which adapts the trained model to unlabeled signals from a new receiver. We first develop a theoretical generalization error bound for the adaptation model. Motivated by the bound, we propose a novel method to solve the cross-receiver RFFI problem, which includes domain alignment and adaptive pseudo-labeling. The former aims at finding a feature space where both domains exhibit similar distributions, effectively reducing the domain discrepancy. Meanwhile, the latter employs a dynamic pseudo-labeling scheme to implicitly transfer the label information from the labeled receiver to the new receiver. Experimental results indicate that the proposed method can effectively mitigate the receiver impact and improve the cross-receiver RFFI performance.


Udio's AI music is my new obsession

PCWorld

My editors yell at me for over-using it. So it's natural that I would latch onto Udio.com, a surprisingly capable AI music generator for everything from death metal to female folk rock. The thing is, I'm not sure how long it will be as great as it is right now. Udio works like an AI art generator. You can either specify a prompt and let Udio do all the heavy lifting, from music to lyrics, or get as detailed as you want.


The Superhero Movie Is Dying. Its Replacement Is Waiting in the Wings.

Slate

For more than a decade, blockbuster comic book adaptations reliably clobbered all competition at the box office. Disney and HBO Max built their streaming strategies around intellectual property from Marvel and DC Comics. The studios turned this pulpy source material into a profusion of interconnected films and series that consistently drove ticket sales and subscriptions--until they didn't. Lately, serious superhero fatigue seems to have set in. Comic book movies regularly tank these days, and not just the ones based on second-string characters like Blue Beetle and Madame Web.