Goto

Collaborating Authors

 Media


MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

arXiv.org Artificial Intelligence

Self-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech. Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is primarily due to the distinctive challenges associated with modelling musical knowledge, particularly its tonal and pitched characteristics of music. To address this research gap, we propose an acoustic Music undERstanding model with large-scale self-supervised Training (MERT), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training. In our exploration, we identified a superior combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance. This combination includes an acoustic teacher based on Residual Vector Quantization - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT). These teachers effectively guide our student model, a BERT-style transformer encoder, to better model music audio. In addition, we introduce an in-batch noise mixture augmentation to enhance the representation robustness. Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters. Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attains state-of-the-art (SOTA) overall scores. The code and models are online: https://github.com/yizhilll/MERT.


Aligning Language Models with Preferences through f-divergence Minimization

arXiv.org Artificial Intelligence

Aligning language models with preferences can be posed as approximating a target distribution representing some desired behavior. Existing approaches differ both in the functional form of the target distribution and the algorithm used to approximate it. For instance, Reinforcement Learning from Human Feedback (RLHF) corresponds to minimizing a reverse KL from an implicit target distribution arising from a KL penalty in the objective. On the other hand, Generative Distributional Control (GDC) has an explicit target distribution and minimizes a forward KL from it using the Distributional Policy Gradient (DPG) algorithm. In this paper, we propose a new approach, f-DPG, which allows the use of any f-divergence to approximate any target distribution that can be evaluated. f-DPG unifies both frameworks (RLHF, GDC) and the approximation methods (DPG, RL with KL penalties). We show the practical benefits of various choices of divergence objectives and demonstrate that there is no universally optimal objective but that different divergences present different alignment and diversity trade-offs. We show that Jensen-Shannon divergence strikes a good balance between these objectives, and frequently outperforms forward KL divergence by a wide margin, leading to significant improvements over prior work. These distinguishing characteristics between divergences persist as the model size increases, highlighting the importance of selecting appropriate divergence objectives.


Self-Adaptive Named Entity Recognition by Retrieving Unstructured Knowledge

arXiv.org Artificial Intelligence

Although named entity recognition (NER) helps us to extract domain-specific entities from text (e.g., artists in the music domain), it is costly to create a large amount of training data or a structured knowledge base to perform accurate NER in the target domain. Here, we propose self-adaptive NER, which retrieves external knowledge from unstructured text to learn the usages of entities that have not been learned well. To retrieve useful knowledge for NER, we design an effective two-stage model that retrieves unstructured knowledge using uncertain entities as queries. Our model predicts the entities in the input and then finds those of which the prediction is not confident. Then, it retrieves knowledge by using these uncertain entities as queries and concatenates the retrieved text to the original input to revise the prediction. Experiments on CrossNER datasets demonstrated that our model outperforms strong baselines by 2.35 points in F1 metric.


Welcome to a World Without Endings

The Atlantic - Technology

Late last month, during yet another inexplicable rebranding exercise, HBO's Max streaming service changed the way it organizes film credits. Rather than separate out discrete production categories for users to peruse, Max's credits lumped writers and directors together under an ominous header, dubbing them "creators." The recategorization enraged writers, filmmakers, and the Directors Guild of America. Within a few hours, Max's parent company, Warner Bros., apologized for the move, calling it "an oversight in the technical transition from HBO Max to Max." The change--made by a company with a market cap that is approaching $30 billion during a contentious writers' strike--felt petty and vindictive to Hollywood professionals.


Indiana Jones 5 gets slammed in reviews - but a new study says poor scores can mean big box office

Daily Mail - Science & tech

Professional movie critics aren't enjoying'Indiana Jones and the Dial of Destiny,' the hotly anticipated, and allegedly final, adventure for Harrison Ford as the whip-cracking archeologist. Pre-release reviews for the picture have lead to a'rotten' 50 percent score at Rotten Tomatoes, based on 46 reviews. And the aggregation site Metacritic gives the new Indy a score of 52/100, based on 24 reviews. But those failing grades could mean that'Indy 5' is shaping up to be a runaway summer sensation -- according to researchers at University of California Davis. Pre-release reviews for'Indiana Jones and the Dial of Destiny' have lead to a'rotten' score of 50% at Rotten Tomatoes and a 52/100 at Metacritic.


Potential memorial designs for Las Vegas massacre unveiled, major step in planning process

FOX News

Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. A series of white angel wings rise up from the earth bathed in a warm glow of light, their sweeping forms creating a long covered pathway surrounded by trees in a possible centerpiece for the memorial to modern America's deadliest mass shooting. It's one of five potential designs unveiled Monday for a permanent monument on the Las Vegas Strip where 58 people were shot and killed and hundreds more injured at a country music festival on Oct. 1, 2017. Two survivors later died from their gunshot wounds.


AI-generated content should be labelled, EU commissioner says

Al Jazeera

Companies deploying AI tools with the ability to generate disinformation, such as ChatGPT and Bard, should label such content as part of their efforts to combat fake news, according to European Commission deputy head Vera Jourova. Unveiled late last year, Microsoft-backed OpenAI's ChatGPT has become the fastest-growing consumer application in history and set off a race among tech companies to bring generative AI products to market. Concerns however are mounting about potential abuse of the technology and the possibility that bad actors and even governments may use it to produce far more disinformation than before. "Signatories who integrate generative AI into their services like Bingchat for Microsoft, Bard for Google should build in necessary safeguards that these services cannot be used by malicious actors to generate disinformation," Jourova told a press conference on Monday. "Signatories who have services with a potential to disseminate AI-generated disinformation should, in turn, put in place technology to recognise such content and clearly label this to users," she said.


'American Pie' icon Don McLean on AI: 'It'll be better than what passes itself off as music today'

FOX News

People in Texas sounded off on AI job displacement, with half of people who spoke to Fox News convinced that the tech will rob them of work. Don McLean, the one-man creative force behind the hit songs "American Pie," "Vincent (Starry, Starry Night)," "And I Love You So," "Castles in the Air," and other songs, albums, tours and projects, shared thoughts about artificial intelligence, music, creativity and authenticity with Fox News Digital in a recent phone interview amid his current "American Pie" 50th anniversary tour. "When you talk about artificial intelligence right now -- I'm not sure what that means at the moment, but clearly it's evolving," he said from California, where he was making several tour stops after returning from concert performances in Australia. "With any technology, you have an inflection point where it takes off," said McLean. "Today, AI has merely presented itself -- but the inflection point hasn't been reached yet. He added, "I also want to say that before a form of artificial intelligence was in use -- and it's been in use for many years -- the tape recorder and the photographic lens were both honest. If you took a picture, that was the way something looked." However, in current times, he said, "you have all this photoshopping and massaging and whatnot, so now the camera lies.


A transparent approach to data representation

arXiv.org Artificial Intelligence

We take inspiration from the non-negative matrix factorization (NMF) problem. In NMF, one large m n In 2006 Netflix released a data set -- roughly 100 million matrix M with non-negative values is factored as a product ratings of 17770 titles, given by 480189 viewers -- of two smaller non-negative matrices R and C of size and posed a challenge: Use this training data to predict m l and l n, respectively (where l m,n). Imagining the ratings in a separate, hidden set of ratings involving the set of ratings as the M matrix, with each row the same movies and viewers. The first to do so with a corresponding to a viewer and each column corresponding root-mean-square prediction error (RMSE) at least 10% to a movie, one can think of each row of R as an lower than that of Netflix's own system would receive a attribute vector for the corresponding viewer.


Hyperbolic Image-Text Representations

arXiv.org Artificial Intelligence

Visual and linguistic concepts naturally organize themselves in a hierarchy, where a textual concept "dog" entails all images that contain dogs. Despite being intuitive, current large-scale vision and language models such as CLIP do not explicitly capture such hierarchy. We propose MERU, a contrastive model that yields hyperbolic representations of images and text. Hyperbolic spaces have suitable geometric properties to embed tree-like data, so MERU can better capture the underlying hierarchy in image-text datasets. Our results show that MERU learns a highly interpretable and structured representation space while being competitive with CLIP's performance on standard multi-modal tasks like image classification and image-text retrieval.