Goto

Collaborating Authors

 Media


Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIP

arXiv.org Artificial Intelligence

Web-crawled datasets have enabled remarkable generalization capabilities in recent image-text models such as CLIP (Contrastive Language-Image pre-training) or Flamingo, but little is known about the dataset creation processes. In this work, we introduce a testbed of six publicly available data sources - YFCC, LAION, Conceptual Captions, WIT, RedCaps, Shutterstock - to investigate how pre-training distributions induce robustness in CLIP. We find that the performance of the pre-training data varies substantially across distribution shifts, with no single data source dominating. Moreover, we systematically study the interactions between these data sources and find that combining multiple sources does not necessarily yield better models, but rather dilutes the robustness of the best individual data source. We complement our empirical findings with theoretical insights from a simple setting, where combining the training data also results in diluted robustness. In addition, our theoretical model provides a candidate explanation for the success of the CLIP-based data filtering technique recently employed in the LAION dataset. Overall our results demonstrate that simply gathering a large amount of data from the web is not the most effective way to build a pre-training dataset for robust generalization, necessitating further study into dataset design. Code is available at https://github.com/mlfoundations/clip_quality_not_quantity.


An Analysis of Classification Approaches for Hit Song Prediction using Engineered Metadata Features with Lyrics and Audio Features

arXiv.org Artificial Intelligence

Hit song prediction, one of the emerging fields in music information retrieval (MIR), remains a considerable challenge. Being able to understand what makes a given song a hit is clearly beneficial to the whole music industry. Previous approaches to hit song prediction have focused on using audio features of a record. This study aims to improve the prediction result of the top 10 hits among Billboard Hot 100 songs using more alternative metadata, including song audio features provided by Spotify, song lyrics, and novel metadata-based features (title topic, popularity continuity and genre class). Five machine learning approaches are applied, including: k-nearest neighbours, Naive Bayes, Random Forest, Logistic Regression and Multilayer Perceptron. Our results show that Random Forest (RF) and Logistic Regression (LR) with all features (including novel features, song audio features and lyrics features) outperforms other models, achieving 89.1% and 87.2% accuracy, and 0.91 and 0.93 AUC, respectively. Our findings also demonstrate the utility of our novel music metadata features, which contributed most to the models' discriminative performance.


Do Multi-Document Summarization Models Synthesize?

arXiv.org Artificial Intelligence

Multi-document summarization entails producing concise synopses of collections of inputs. For some applications, the synopsis should accurately \emph{synthesize} inputs with respect to a key property or aspect. For example, a synopsis of film reviews all written about a particular movie should reflect the average critic consensus. As a more consequential example, consider narrative summaries that accompany biomedical \emph{systematic reviews} of clinical trial results. These narratives should fairly summarize the potentially conflicting results from individual trials. In this paper we ask: To what extent do modern multi-document summarization models implicitly perform this type of synthesis? To assess this we perform a suite of experiments that probe the degree to which conditional generation models trained for summarization using standard methods yield outputs that appropriately synthesize inputs. We find that existing models do partially perform synthesis, but do so imperfectly. In particular, they are over-sensitive to changes in input ordering and under-sensitive to changes in input compositions (e.g., the ratio of positive to negative movie reviews). We propose a simple, general method for improving model synthesis capabilities by generating an explicitly diverse set of candidate outputs, and then selecting from these the string best aligned with the expected aggregate measure for the inputs, or \emph{abstaining} when the model produces no good candidate. This approach improves model synthesis performance. We hope highlighting the need for synthesis (in some summarization settings), motivates further research into multi-document summarization methods and learning objectives that explicitly account for the need to synthesize.


Automated Sentiment and Hate Speech Analysis of Facebook Data by Employing Multilingual Transformer Models

arXiv.org Artificial Intelligence

In recent years, there has been a heightened consensus - both within academia and in the public discourse - that Social Media Platforms (SMPs), amplify the spread of hateful and negative sentiment content. Researchers have identified how hateful content, political propaganda, and targeted messaging contributed to real-world harms including insurrections against democratically elected governments, genocide, and breakdown of social cohesion due to heightened negative discourse towards certain communities in parts of the world. To counter these issues, SMPs have created semi-automated systems that can help identify toxic speech. In this paper we analyse the statistical distribution of hateful and negative sentiment contents within a representative Facebook dataset (n= 604,703) scrapped through 648 public Facebook pages which identify themselves as proponents (and followers) of far-right Hindutva actors. These pages were identified manually using keyword searches on Facebook and on CrowdTangleand classified as far-right Hindutva pages based on page names, page descriptions, and discourses shared on these pages. We employ state-of-the-art, open-source XLM-T multilingual transformer-based language models to perform sentiment and hate speech analysis of the textual contents shared on these pages over a period of 5.5 years. The result shows the statistical distributions of the predicted sentiment and the hate speech labels; top actors, and top page categories. We further discuss the benchmark performances and limitations of these pre-trained language models.


The Efficacy of Self-Supervised Speech Models for Audio Representations

arXiv.org Artificial Intelligence

Self-supervised learning (SSL) speech models, which can serve as powerful upstream models to extract meaningful speech representations, have achieved unprecedented success in speech representation learning. However, their effectiveness on non-speech datasets is relatively less explored. In this work, we propose an ensemble framework, with a combination of ensemble techniques, to fuse SSL speech models' embeddings. Extensive experiments on speech and non-speech audio datasets are conducted to investigate the representation abilities of our ensemble method and its single constituent model. Ablation studies are carried out to evaluate the performances of different ensemble techniques, such as feature averaging and concatenation. All experiments are conducted during NeurIPS 2021 HEAR Challenge as a standard evaluation pipeline provided by competition officials. Results demonstrate SSL speech models' strong abilities on various non-speech tasks, while we also note that they fail to deal with fine-grained music tasks, such as pitch classification and note onset detection. In addition, feature ensemble is shown to have great potential on producing more holistic representations, as our proposed framework generally surpasses state-of-the-art SSL speech/audio models and has superior performance on various datasets compared with other teams in HEAR Challenge.


Archive TimeLine Summarization (ATLS): Conceptual Framework for Timeline Generation over Historical Document Collections

arXiv.org Artificial Intelligence

Archive collections are nowadays mostly available through search engines interfaces, which allow a user to retrieve documents by issuing queries. The study of these collections may be, however, impaired by some aspects of search engines, such as the overwhelming number of documents returned or the lack of contextual knowledge provided. New methods that could work independently or in combination with search engines are then required to access these collections. In this position paper, we propose to extend TimeLine Summarization (TLS) methods on archive collections to assist in their studies. We provide an overview of existing TLS methods and we describe a conceptual framework for an Archive TimeLine Summarization (ATLS) system, which aims to generate informative, readable and interpretable timelines.


How Artificial Intelligence (AI) will impact Hollywood

#artificialintelligence

AI is rapidly changing the way Hollywood functions. It revolutionizes how stories are told, how movies are made, how audiences engage with content, and more. AI has the potential to disrupt the entire movie industry, from the way producers develop scripts to the way audiences consume content. AI is already being used to help filmmakers create more engaging stories. AI-powered screenwriting tools are being used to help writers generate ideas and structure their stories.


Google Has AI, Can Make Music from Text

#artificialintelligence

In its research paper, the AI, called MusicLM, claims to be a model that can create high-fidelity music from text descriptions such as'a soothing violin melody with distorted guitar riffs'. "We demonstrated that MusicLM can be conditioned using text and melodies where they can change the whistling and humming melodies according to the style described in the text captions," Google researchers said in the paper, as quoted from Insider, Monday (1/30/2023) . According to the study, users could enter text descriptions such as'nice jazz with saxophone solo and soloist', or '90s-style Techno Berlin with low bass' and then receive music similar to that text query. Currently users cannot use MusicLM to create music. But you can listen to his music on the following Google Github page. To compete with ChatGPT, Google will release 20 new products, including a new version of Google Search with an AI chatbot feature.


+55 Best AI Tools To Improve Your Writing Skills - LOST OFFER

#artificialintelligence

In today's digital age, AI has become an increasingly useful tool for a variety of tasks, including writing. Whether you're a novelist, journalist, or student, AI can really help you improve your writing skills. In this article, we will explore some of the ways in which you can use AI to enhance your writing skills and improve the quality of your work. Also, We have made you a list of some Free AI Writing Tools to get you started right away with some tips to keep in mind while using these AI Writing Tools. The term "AI" has become increasingly prevalent in both home and office settings. You may have heard your friends or family members talking about how they are using AI to improve their writing skills, or you may have seen it mentioned in an article or on social media. But what exactly is AI?. In simple terms, AI is a computer program that is designed to learn and improve its performance over time.


ChatGPT's answer on which is the most reliable finance news portal in India. Watch

#artificialintelligence

When asked ChatGPT which is the most reliable financial news portal in India and here's what the artificial intelligence (AI) chatbot had to say: "Moneycontrol is considered to be one of the leading finance news portals in India. It provides financial news, analysis, and information on stocks, mutual funds, and other financial instruments. It also provides tools such as stock screener, portfolio tracker and alerts, and financial calculators to help users make informed decisions. The Budget 2023-24 will be broadcast live on Lok Sabha TV along with other news outlets. It can also be viewed live on Moneycontrol.com.