Goto

Collaborating Authors

 Deep Learning


Spectrogram Feature Losses for Music Source Separation

arXiv.org Machine Learning

Abstract--In this paper we study deep learning-based music source separation, and explore using an alternative loss to the standard spectrogram pixel-level L2 loss for model training. Our main contribution is in demonstrating that adding a highlevel featureloss term, extracted from the spectrograms using a VGG net, can improve separation quality visa-vis a pure pixel-level loss. We show this improvement in the context of the MMDenseNet, a State-of-the-Art deep learning model for this task, for the extraction of drums and vocal sounds from songs in the musdb18 database, covering a broad range of western music genres. We believe that this finding can be generalized and applied to broader machine learning-based systems in the audio domain. I. INTRODUCTION Music source separation is a problem that has been studied for a few decades now: given an audio track with several instruments mixed together (a regular MP3 file, for example), how can it be separated into its component instruments? The obvious application of this problem is in music production - creating karaoke tracks, highlighting select instruments in an audio playback, etc.


Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

arXiv.org Machine Learning

Transformer networks have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. As a solution, we propose a novel neural architecture, Transformer-XL, that enables Transformer to learn dependency beyond a fixed length without disrupting temporal coherence. Concretely, it consists of a segment-level recurrence mechanism and a novel positional encoding scheme. Our method not only enables capturing longer-term dependency, but also resolves the problem of context fragmentation. As a result, Transformer-XL learns dependency that is about 80% longer than RNNs and 450% longer than vanilla Transformers, achieves better performance on both short and long sequences, and is up to 1,800 times faster than vanilla Transformer during evaluation. Language modeling is among the important problems that require modeling long-term dependency, with successful applications such as unsupervised pretraining (Dai & Le, 2015; Peters et al., 2018; Radford et al., 2018; Devlin et al., 2018). However, it has been a challenge to equip neural networks with the capability to model long-term dependency in sequential data. Recurrent neural networks (RNNs), in particular Long Short-Term Memory (LSTM) networks (Hochreiter & Schmidhuber, 1997), have been a standard solution to language modeling and obtained strong results on multiple benchmarks. Despite the wide adaption, RNNs are difficult to optimize due to gradient vanishing and explosion (Hochreiter et al., 2001), and the introduction of gating in LSTMs and the gradient clipping technique (Graves, 2013; Pascanu et al., 2012) might not be sufficient to fully address this issue. Empirically, previous work has found that LSTM language models use 200 context words on average (Khandelwal et al., 2018), indicating room for further improvement. On the other hand, the direct connections between long-distance word pairs baked in attention mechanisms might ease optimization and enable the learning of long-term dependency (Bahdanau et al., 2014; V aswani et al., 2017).


A New Human Ancestor Has Been Discovered Thanks To Artificial Intelligence

#artificialintelligence

It's a well-known fact that there was a lot of interspecies mingling back in the day. Modern humans have fragments of DNA from our ancient relatives, the Neanderthals and the Denisovans โ€“ and now, a third, previously unknown, mystery species. An international team of researchers have examined human DNA using deep learning algorithms to analyze genetic clues to human evolution for the very first time. The results are published in the journal Nature Communications. It's long been suspected that, further to the Neanderthals and Denisovans, people of Asian descent have a third ancestor that interbred with ancient humans.


Deep Learning Finds Fake News with 97% Accuracy

#artificialintelligence

That means the pooling layer computes a feature vector of size 128 which is passed into dense layers of the feedforward network as we mentioned above. The overall structure of the DNN can be understood as a preprocessor defined in the first part that is being trained to map text sequences into feature vectors in such a way that the weights of the second part can be trained to obtain optimal classification results from the overall network. More details on the implementation and text preprocessing can be found in my GitHub repository for this project. I trained this network for 10 epochs with a batch size of 128 using an 80-20 training/hold-out set. A couple of notes on additional parameters: The vast majority of documents in this collection is of length 5000 or less. So for the maximum input sequence length for the DNN I chose 5000 words. There are roughly 100,000 unique words in this collection of documents. I arbitrarily limited the dictionary that the DNN can learn to 25% of that: 25,000 words. Finally, for the embedding dimension, I chose 300 simply because that is the default embedding dimension for both word2vec and GloVe.


Building My First Deep Learning Machine

#artificialintelligence

First of all, myth busted: the 1080 Ti can run minesweeper effortlessly. The machine did restart itself once for no obvious reasons after the proprietary GPU driver was installed. Back to the topicโ€ฆ Here is some R code for fitting a "wide and deep" classification model with Tensorflow and Tensorflow Estimators API. The model is fundamentally a direct combination of a linear model and a DNN model. The synthetic data has 1 million observations, 100 features (20 being useful) and is generated by my R package msaenet.


Game Theory for Data Scientists โ€“ Towards Data Science

#artificialintelligence

Games are playing a key role in the evolution of artificial intelligence(AI). For starters, game environments are becoming a popular training mechanism in areas such as reinforcement learning or imitation learning. In theory, any multi-agent AI system can be subjected to gamified interactions between its participants. The branch of mathematics that formulates the principles of games is known as game theory. In the context of artificial intelligence(AI) and deep learning systems, game theory is essential to enable some of the key capabilities required in multi-agent environments in which different AI programs need to interact or compete in order to accomplish a goal.


A deep learning-based method to detect cyberbullying on Twitter

#artificialintelligence

Researchers at King Saud University, in Saudi Arabia, have developed a new approach to detect cyberbullying on Twitter using deep learning called OCDD. In contrast with other deep-learning approaches, which extract features from tweets and feed them to a classifier, their method represents a tweet as a set of word vectors. In recent years, cyberbullying on social media has become a huge and widely discussed issue. Cyberbullying entails the use of online communication channels to bully other users by sending intimidating, threatening or abusive messages. This can have psychological and sometimes life-threatening consequences for the victims.


Global Artificial Intelligence (AI) Industry

#artificialintelligence

Germany Market Analysis Table 35: German Recent Past, Current & Future Analysis for Artificial Intelligence Analyzed with Annual Revenue Figures in US$ Million for Years 2015 through 2024 (includes corresponding Graph/Chart) 9.4.3 Italy Market Analysis Table 36: Italian Recent Past, Current & Future Analysis for Artificial Intelligence Analyzed with Annual Revenue Figures in US$ Million for Years 2015 through 2024 (includes corresponding Graph/Chart) 9.4.4


Pix2Story: Neural storyteller which creates machine-generated story in several literature genre Blog Microsoft Azure

#artificialintelligence

Storytelling is at the heart of human nature. We were storytellers long before we were able to write, we shared our values and created our societies mostly through oral storytelling. Then, we managed to find the way to record and share our stories, and certainly more advanced ways to broadly share our stories; from Gutenberg's printing press to television, and the internet. Writing stories is not easy, especially if one must write a story just by looking at a picture in different literary genres. Natural Language Processing (NLP) is a field that is driving a revolution in the computer-human interaction.


Artificial intelligence can now predict if someone will die in the next 5 years

#artificialintelligence

This AI will tell people when theyre likely to die -- and thats a good thing. Thats because scientists from the University of Adelaide in Australia have used deep learning technology to analyze the computerized tomography (CT) scans of patient organs, in what could one day serve as an early warning system to catch heart disease, cancer, and other diseases early so that intervention can take place. Using a dataset of historical CT scans, and excluding other predictive factors like age, the system developed by the team was able to predict whether patients would die within five years around 70 percent of the time. The work was described in an article published in the journal Scientific Reports. The goal of the research isn't really to predict death, but to produce a more accurate measurement of health, Dr. Luke Oakden-Rayner, a researcher on the project, told Digital Trends.