Media
GAFX: A General Audio Feature eXtractor
Bu, Zhaoyang, Zhang, Hanhaodi, Zhu, Xiaohu
Most machine learning models for audio tasks are dealing with a handcrafted feature, the spectrogram. However, it is still unknown whether the spectrogram could be replaced with deep learning based features. In this paper, we answer this question by comparing the different learnable neural networks extracting features with a successful spectrogram model and proposed a General Audio Feature eXtractor (GAFX) based on a dual U-Net (GAFX-U), ResNet (GAFX-R), and Attention (GAFX-A) modules. We design experiments to evaluate this model on the music genre classification task on the GTZAN dataset and perform a detailed ablation study of different configurations of our framework and our model GAFX-U, following the Audio Spectrogram Transformer (AST) classifier achieves competitive performance.
Differentiable Time-Frequency Scattering on GPU
Muradeli, John, Vahidi, Cyrus, Wang, Changhong, Han, Han, Lostanlen, Vincent, Lagrange, Mathieu, Fazekas, George
Joint time-frequency scattering (JTFS) is a convolutional operator in the time-frequency domain which extracts spectrotemporal modulations at various rates and scales. It offers an idealized model of spectrotemporal receptive fields (STRF) in the primary auditory cortex, and thus may serve as a biological plausible surrogate for human perceptual judgments at the scale of isolated audio events. Yet, prior implementations of JTFS and STRF have remained outside of the standard toolkit of perceptual similarity measures and evaluation methods for audio generation. We trace this issue down to three limitations: differentiability, speed, and flexibility. In this paper, we present an implementation of time-frequency scattering in Python. Unlike prior implementations, ours accommodates NumPy, PyTorch, and TensorFlow as backends and is thus portable on both CPU and GPU. We demonstrate the usefulness of JTFS via three applications: unsupervised manifold learning of spectrotemporal modulations, supervised classification of musical instruments, and texture resynthesis of bioacoustic sounds.
QuoteKG: A Multilingual Knowledge Graph of Quotes
Kuculo, Tin, Gottschalk, Simon, Demidova, Elena
Quotes of public figures can mark turning points in history. A quote can explain its originator's actions, foreshadowing political or personal decisions and revealing character traits. Impactful quotes cross language barriers and influence the general population's reaction to specific stances, always facing the risk of being misattributed or taken out of context. The provision of a cross-lingual knowledge graph of quotes that establishes the authenticity of quotes and their contexts is of great importance to allow the exploration of the lives of important people as well as topics from the perspective of what was actually said. In this paper, we present QuoteKG, the first multilingual knowledge graph of quotes. We propose the QuoteKG creation pipeline that extracts quotes from Wikiquote, a free and collaboratively created collection of quotes in many languages, and aligns different mentions of the same quote. QuoteKG includes nearly one million quotes in $55$ languages, said by more than $69,000$ people of public interest across a wide range of topics. QuoteKG is publicly available and can be accessed via a SPARQL endpoint.
A Hybrid Recommender System for Recommending Smartphones to Prospective Customers
Biswas, Pratik K., Liu, Songlin
Recommender Systems are a subclass of machine learning systems that employ sophisticated information filtering strategies to reduce the search time and suggest the most relevant items to any particular user. Hybrid recommender systems combine multiple recommendation strategies in different ways to benefit from their complementary advantages. Some hybrid recommender systems have combined collaborative filtering and content-based approaches to build systems that are more robust. In this paper, we propose a hybrid recommender system, which combines Alternating Least Squares (ALS) based collaborative filtering with deep learning to enhance recommendation performance as well as overcome the limitations associated with the collaborative filtering approach, especially concerning its cold start problem. In essence, we use the outputs from ALS (collaborative filtering) to influence the recommendations from a Deep Neural Network (DNN), which combines characteristic, contextual, structural and sequential information, in a big data processing framework. We have conducted several experiments in testing the efficacy of the proposed hybrid architecture in recommending smartphones to prospective customers and compared its performance with other open-source recommenders. The results have shown that the proposed system has outperformed several existing hybrid recommender systems.
Comprehensive Analysis of the Object Detection Pipeline on UAVs
Varga, Leon Amadeus, Koch, Sebastian, Zell, Andreas
An object detection pipeline comprises a camera that captures the scene and an object detector that processes these images. The quality of the images directly affects the performance of the object detector. Many works nowadays focus either on improving the image quality or improving the object detection models independently, but neglect the importance of joint optimization of the two subsystems. The goal of this paper is to tune the detection throughput and accuracy of existing object detectors in the remote sensing scenario by focusing on optimizing the input images tailored to the object detector. To achieve this, we empirically analyze the influence of two selected camera calibration parameters (camera distortion correction and gamma correction) and five image parameters (quantization, compression, resolution, color model, additional channels) for these applications. For our experiments, we utilize three UAV data sets from different domains and a mixture of large and small state-of-the-art object detector models to provide an extensive evaluation of the influence of the pipeline parameters. Finally, we realize an object detection pipeline prototype on an embedded platform for an UAV and give a best practice recommendation for building object detection pipelines based on our findings. We show that not all parameters have an equal impact on detection accuracy and data throughput, and that by using a suitable compromise between parameters we are able to achieve higher detection accuracy for lightweight object detection models, while keeping the same data throughput.
Mimetic Models: Ethical Implications of AI that Acts Like You
McIlroy-Young, Reid, Kleinberg, Jon, Sen, Siddhartha, Barocas, Solon, Anderson, Ashton
An emerging theme in artificial intelligence research is the creation of models to simulate the decisions and behavior of specific people, in domains including game-playing, text generation, and artistic expression. These models go beyond earlier approaches in the way they are tailored to individuals, and the way they are designed for interaction rather than simply the reproduction of fixed, pre-computed behaviors. We refer to these as mimetic models, and in this paper we develop a framework for characterizing the ethical and social issues raised by their growing availability. Our framework includes a number of distinct scenarios for the use of such models, and considers the impacts on a range of different participants, including the target being modeled, the operator who deploys the model, and the entities that interact with it.
Tools to Use When Building Sentiment Analyzer
Originally published on Towards AI the World's Leading AI and Technology News and Media Company. If you are building an AI-related product or service, we invite you to consider becoming an AI sponsor. At Towards AI, we help scale AI and technology startups. Let us help you unleash your technology to the masses. Sentiment Analysis is a powerful tool to use when trying to understand how to test your website.
Role of AI in Software Development - Analytics Vidhya
This article was published as a part of the Data Science Blogathon. In the 21st Century, we were introduced to the concept of automation. Data and machine intelligence are powering this automation. Data simply means a collection of values. In 2022, the per day Data generation stats stand at a staggering 2.5 quintillion bytes, the unit conversion of which to Gigabytes becomes a head-spinning task.
How Artificial Intelligence is Regulating Live Video Streams
It's possible that you may have already come across Artificial Intelligence (AI) at some point in your life without even realizing it. For example, Facebook, Twitter, and Google all use AI to ensure that users have a seamless experience on their platforms, whether by automatically tagging friends in photos or providing results based on previous searches. These uses of AI are relatively simple and only involve one part of the technology โ Machine Learning (ML). Fundamentally, ML is becoming more prevalent, but what about its big brother and sister, Deep Learning (DL), and narrow Artificial Intelligence? How can these potentially create streaming services that we will never want to live without?
This Band Wrote the Best Legend of Zelda Song of 2022
Horse Jumper of Love is a rock band from Boston that makes the kind of music you might want playing in your hyperbaric chamber if you were stuck in there for a while and really wanted to lean into the experience. One of their best tracks, 2019's "DIRT," is built around a piercingly plonking guitar riff and the phrase "And there is dirt and there is juice / and I am mixing up the two." I don't know what it means. I'm not sure I'm supposed to. The strange slowcore formulations of their new album, Natural Part, is full of similarly perplexing songwriting.