Goto

Collaborating Authors

 Media


One Billion Audio Sounds from GPU-enabled Modular Synthesis

arXiv.org Artificial Intelligence

We release synth1B1, a multi-modal audio corpus consisting of 1 billion 4-second synthesized sounds, which is 100x larger than any audio dataset in the literature. Each sound is paired with the corresponding latent parameters used to generate it. synth1B1 samples are deterministically generated on-the-fly 16200x faster than real-time (714MHz) on a single GPU using torchsynth (https://github.com/torchsynth/torchsynth), an open-source modular synthesizer we release. Additionally, we release two new audio datasets: FM synth timbre (https://zenodo.org/record/4677102) and subtractive synth pitch (https://zenodo.org/record/4677097). Using these datasets, we demonstrate new rank-based synthesizer-motivated evaluation criteria for existing audio representations. Finally, we propose novel approaches to synthesizer hyperparameter optimization, and demonstrate how perceptually-correlated auditory distances could enable new applications in synthesizer design.


[R] Google-Workshop: Conceptual Understanding of Deep Learning, May 17. Join Us.

#artificialintelligence

Please join us for a virtual Google workshop on "Conceptual Understanding of Deep Learning" When: May 17th 9am-4pm PST. Goal: How does the Brain/Mind (perhaps even an artificial one) work at an algorithmic level? While deep learning has produced tremendous technological strides in recent decades, there is an unsettling feeling of a lack of "conceptual" understanding of why it works and to what extent it will work in the current form. The goal of the workshop is to bring together theorists and practitioners to develop an understanding of the right algorithmic view of deep learning, characterizing the class of functions that can be learned, coming up with the right learning architecture that may (provably) learn multiple functions, concepts and remember them over time as humans do, theoretical understanding of language, logic, RL, meta learning and lifelong learning. The speakers and panelists include Turing award winners Geoffrey Hinton, Leslie Valiant, and Godel Prize winner Christos Papadimitriou (full-details).


[D] Do I practice on one area of ML or is it better to practice on all algorithms?

#artificialintelligence

I am an intermediate in ML and I know the theory of several supervised and unsupervised ML algorithms as well as deep learning. However, I lack huge amount of practical skills which is what I am working on right now. However, I am lost as to what to practice exactly. I am mostly interested in deep learning but I am seeing how essential it is to know how to implement other algorithms as well like random forests, SVM.. etc. Do I practice DL and other ML algorithms simultaneously (as mastering one area is almost impossible and is done across many years of experience) or should I first focus on one area (say computer vision) and then move on to the rest?


The insider pro trick to find any photo on your phone in seconds

FOX News

Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. Our phones are jam-packed with photos. Pick 25 at random, and I bet only a handful are decent photos you want to keep around. Duplicates and the shot right before the good one make up a lot of that junk.


15 tech tips you won't find in a user manual

FOX News

Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. Most gadgets don't come with a user manual that spells out every single feature. We learn them by doing, when someone spills the beans, or asking, "How'd you do that?" For example, no one thinks to dive into a new router's settings.


AI skills are a problem. AutoML can help

#artificialintelligence

The biggest barrier to enterprise success with AI is difficulty finding people with the requisite skills.



Deep Probabilistic Graphical Modeling

arXiv.org Machine Learning

Probabilistic graphical modeling (PGM) provides a framework for formulating an interpretable generative process of data and expressing uncertainty about unknowns, but it lacks flexibility. Deep learning (DL) is an alternative framework for learning from data that has achieved great empirical success in recent years. DL offers great flexibility, but it lacks the interpretability and calibration of PGM. This thesis develops deep probabilistic graphical modeling (DPGM.) DPGM consists in leveraging DL to make PGM more flexible. DPGM brings about new methods for learning from data that exhibit the advantages of both PGM and DL. We use DL within PGM to build flexible models endowed with an interpretable latent structure. One model class we develop extends exponential family PCA using neural networks to improve predictive performance while enforcing the interpretability of the latent factors. Another model class we introduce enables accounting for long-term dependencies when modeling sequential data, which is a challenge when using purely DL or PGM approaches. Finally, DPGM successfully solves several outstanding problems of probabilistic topic models, a widely used family of models in PGM. DPGM also brings about new algorithms for learning with complex data. We develop reweighted expectation maximization, an algorithm that unifies several existing maximum likelihood-based algorithms for learning models parameterized by neural networks. This unifying view is made possible using expectation maximization, a canonical inference algorithm in PGM. We also develop entropy-regularized adversarial learning, a learning paradigm that deviates from the traditional maximum likelihood approach used in PGM. From the DL perspective, entropy-regularized adversarial learning provides a solution to the long-standing mode collapse problem of generative adversarial networks, a widely used DL approach.


Music Embedding: A Tool for Incorporating Music Theory into Computational Music Applications

arXiv.org Artificial Intelligence

Advancements in the digital technologies have enabled researchers to develop a variety of Computational Music applications. Such applications are required to capture, process, and generate data related to music. Therefore, it is important to digitally represent music in a music theoretic and concise manner. Existing approaches for representing music are ineffective in terms of utilizing music theory. In this paper, we address the disjoint of music theory and computational music by developing an opensource representation tool based on music theory. Through the wide range of use cases, we run an analysis on the classical music pieces to show the usefulness of the developed music embedding.


Kazuo Ishiguro writes of artificial intelligence and human hearts in 'Klara and the Sun'

#artificialintelligence

Klara, the narrator of the new novel by Kazuo Ishiguro, isn't human, but understanding humans is her mission. In Klara and the Sun, the reader follows her in that mission, in a world that seems like our own in a none too distant future. Ishiguro, who was born in Japan but has lived most of his life in England, has written seven previous novels, including the Booker Prize-winning The Remains of the Day, as well as short fiction, song lyrics and screenplays. Klara and the Sun is his first novel since he received the Nobel Prize for literature in 2017. It underscores how well he deserved that prize, in its beautiful craft and prose and in its tender but unflinching sense of the human heart.