Goto

Collaborating Authors

 Media


Music Instrument Classification Reprogrammed

arXiv.org Artificial Intelligence

The performance of approaches to Music Instrument Classification, a popular task in Music Information Retrieval, is often impacted and limited by the lack of availability of annotated data for training. We propose to address this issue with "reprogramming," a technique that utilizes pre-trained deep and complex neural networks originally targeting a different task by modifying and mapping both the input and output of the pre-trained model. We demonstrate that reprogramming can effectively leverage the power of the representation learned for a different task and that the resulting reprogrammed system can perform on par or even outperform state-of-the-art systems at a fraction of training parameters. Our results, therefore, indicate that reprogramming is a promising technique potentially applicable to other tasks impeded by data scarcity.


RoMQA: A Benchmark for Robust, Multi-evidence, Multi-answer Question Answering

arXiv.org Artificial Intelligence

We introduce RoMQA, the first benchmark for robust, multi-evidence, multi-answer question answering (QA). RoMQA contains clusters of questions that are derived from related constraints mined from the Wikidata knowledge graph. RoMQA evaluates robustness of QA models to varying constraints by measuring worst-case performance within each question cluster. Compared to prior QA datasets, RoMQA has more human-written questions that require reasoning over more evidence text and have, on average, many more correct answers. In addition, human annotators rate RoMQA questions as more natural or likely to be asked by people. We evaluate state-of-the-art large language models in zero-shot, few-shot, and fine-tuning settings, and find that RoMQA is challenging: zero-shot and few-shot models perform similarly to naive baselines, while supervised retrieval methods perform well below gold evidence upper bounds. Moreover, existing models are not robust to variations in question constraints, but can be made more robust by tuning on clusters of related questions. Our results show that RoMQA is a challenging benchmark for large language models, and provides a quantifiable test to build more robust QA methods.


AI & Future Of The Lens & Screen Arts

#artificialintelligence

The growing popularity of tools like DALL-E & MidJourney illustrates how artificial intelligence (AI) is quickly transforming image-making. Digital culture theorist & artist Lev Manovich will be in conversation with media scholar & writer Natasha Chuk to share examples & discuss these developments. They will address the emerging aesthetics & creative practices of AI photography & what these changes mean for the future of photography & the lens-based arts. This event is the first of a series of events & discussions exploring the relationship between AI & the lens & screen arts hosted by the MFA Photography, Video & Related Media Department at the School of Visual Arts. Stay tuned for future events in early 2023.


'I lie in the bath, imagining that I am wandering the Rialto in Venice': my obsession with Duolingo

The Guardian

This morning, before checking in on my young son or making a coffee, I opened the Duolingo app on my phone and translated "They love smelling meat" into Italian. I've been starting my days like this for a few months now: wake up, wash face, grapple with the gerund. I usually spend between 10 and 20 minutes on it while the kettle boils or I load CBeebies or write some emails. Duolingo is a language learning app and pretty simple to use. After you've chosen which language you want to learn, you are presented with about 100 skill-sets divided by scenario or grammar (grocery shopping, the future tense and so on).


VIZIO Unveils Reimagined User Experience and Branding for WatchFree+

#artificialintelligence

VIZIO has unveiled an upgraded design and user experience for WatchFree, VIZIO's free streaming service that comes built into millions of VIZIO Smart TVs. This latest update brings a new look and feel, intuitive Electronic Program Guide (EPG), faster and easier navigation, and personalization features to the free streaming service. WatchFree has grown to include more than 260 free channels and 6,000 titles on demand, offering an ever-expanding library of movies, TV shows, news, sports, music, and programs. "We are proud of the growth we have seen across our free streaming service" The redesigned WatchFree reflects VIZIO's commitment to its users, with features centered around an enhanced user experience and access to fan-favorite programming. The WatchFree Guide now displays up-coming content segmented in 30-minute intervals, allowing viewers to explore the vast collection of WatchFree content in a familiar format.


AI Drew This Gorgeous Comic Series, But You'd Never Know It

#artificialintelligence

You might expect a comic book series featuring art generated entirely by artificial intelligence technology to be full of surreal images that have you tilting your head trying to grasp what kind of sense-shifting madness you're looking at. Not so with the images in The Bestiary Chronicles, a free, three-part comics series from Campfire Entertainment, an award-winning New York-based production house focused on creative storytelling. In The Lesson, a teacher tells students about the monsters that ruined their planet. The team behind the comic used the phrase "Hitchcock Blonde" to describe the story's heroine to AI art-generation tool Midjourney, "and more often than not she came out looking like Grace Kelly," says writer Steve Coulson. The visuals in the trilogy -- believed to be the first comics series made with AI-assisted art -- are stunning.


10 things to try with your new Google Home smart speaker

#artificialintelligence

Did you miss a session from GamesBeat Summit Next 2022? All sessions are now available for viewing in our on-demand library. Click here to start watching. With Google Assistant inside and conversational AI, these speakers can do a great range of things. Here's 10 worth trying, drawn from VentureBeat coverage over the course of the past year. Before getting into the more dynamic features Google Assistant provides through Home smart speakers, start with the most popular ways people use speakers with intelligent assistants.


Learning to Answer Multilingual and Code-Mixed Questions

arXiv.org Artificial Intelligence

Question-answering (QA) that comes naturally to humans is a critical component in seamless human-computer interaction. It has emerged as one of the most convenient and natural methods to interact with the web and is especially desirable in voice-controlled environments. Despite being one of the oldest research areas, the current QA system faces the critical challenge of handling multilingual queries. To build an Artificial Intelligent (AI) agent that can serve multilingual end users, a QA system is required to be language versatile and tailored to suit the multilingual environment. Recent advances in QA models have enabled surpassing human performance primarily due to the availability of a sizable amount of high-quality datasets. However, the majority of such annotated datasets are expensive to create and are only confined to the English language, making it challenging to acknowledge progress in foreign languages. Therefore, to measure a similar improvement in the multilingual QA system, it is necessary to invest in high-quality multilingual evaluation benchmarks. In this dissertation, we focus on advancing QA techniques for handling end-user queries in multilingual environments. This dissertation consists of two parts. In the first part, we explore multilingualism and a new dimension of multilingualism referred to as code-mixing. Second, we propose a technique to solve the task of multi-hop question generation by exploiting multiple documents. Experiments show our models achieve state-of-the-art performance on answer extraction, ranking, and generation tasks on multiple domains of MQA, VQA, and language generation. The proposed techniques are generic and can be widely used in various domains and languages to advance QA systems.


Exploiting Device and Audio Data to Tag Music with User-Aware Listening Contexts

arXiv.org Artificial Intelligence

As music has become more available especially on music streaming platforms, people have started to have distinct preferences to fit to their varying listening situations, also known as context. Hence, there has been a growing interest in considering the user's situation when recommending music to users. Previous works have proposed user-aware autotaggers to infer situation-related tags from music content and user's global listening preferences. However, in a practical music retrieval system, the autotagger could be only used by assuming that the context class is explicitly provided by the user. In this work, for designing a fully automatised music retrieval system, we propose to disambiguate the user's listening information from their stream data. Namely, we propose a system which can generate a situational playlist for a user at a certain time 1) by leveraging user-aware music autotaggers, and 2) by automatically inferring the user's situation from stream data (e.g. device, network) and user's general profile information (e.g. age). Experiments show that such a context-aware personalized music retrieval system is feasible, but the performance decreases in the case of new users, new tracks or when the number of context classes increases.


YM2413-MDB: A Multi-Instrumental FM Video Game Music Dataset with Emotion Annotations

arXiv.org Artificial Intelligence

Existing multi-instrumental datasets tend to be biased toward pop and classical music. In addition, they generally lack high-level annotations such as emotion tags. In this paper, we propose YM2413-MDB, an 80s FM video game music dataset with multi-label emotion annotations. It includes 669 audio and MIDI files of music from Sega and MSX PC games in the 80s using YM2413, a programmable sound generator based on FM. The collected game music is arranged with a subset of 15 monophonic instruments and one drum instrument. They were converted from binary commands of the YM2413 sound chip. Each song was labeled with 19 emotion tags by two annotators and validated by three verifiers to obtain refined tags. We provide the baseline models and results for emotion recognition and emotion-conditioned symbolic music generation using YM2413-MDB.