Media
One Billion Audio Sounds from GPU-enabled Modular Synthesis
Turian, Joseph, Shier, Jordie, Tzanetakis, George, McNally, Kirk, Henry, Max
We release synth1B1, a multi-modal audio corpus consisting of 1 billion 4-second synthesized sounds, which is 100x larger than any audio dataset in the literature. Each sound is paired with the corresponding latent parameters used to generate it. synth1B1 samples are deterministically generated on-the-fly 16200x faster than real-time (714MHz) on a single GPU using torchsynth (https://github.com/torchsynth/torchsynth), an open-source modular synthesizer we release. Additionally, we release two new audio datasets: FM synth timbre (https://zenodo.org/record/4677102) and subtractive synth pitch (https://zenodo.org/record/4677097). Using these datasets, we demonstrate new rank-based synthesizer-motivated evaluation criteria for existing audio representations. Finally, we propose novel approaches to synthesizer hyperparameter optimization, and demonstrate how perceptually-correlated auditory distances could enable new applications in synthesizer design.
[R] Google-Workshop: Conceptual Understanding of Deep Learning, May 17. Join Us.
Please join us for a virtual Google workshop on "Conceptual Understanding of Deep Learning" When: May 17th 9am-4pm PST. Goal: How does the Brain/Mind (perhaps even an artificial one) work at an algorithmic level? While deep learning has produced tremendous technological strides in recent decades, there is an unsettling feeling of a lack of "conceptual" understanding of why it works and to what extent it will work in the current form. The goal of the workshop is to bring together theorists and practitioners to develop an understanding of the right algorithmic view of deep learning, characterizing the class of functions that can be learned, coming up with the right learning architecture that may (provably) learn multiple functions, concepts and remember them over time as humans do, theoretical understanding of language, logic, RL, meta learning and lifelong learning. The speakers and panelists include Turing award winners Geoffrey Hinton, Leslie Valiant, and Godel Prize winner Christos Papadimitriou (full-details).
[D] Do I practice on one area of ML or is it better to practice on all algorithms?
I am an intermediate in ML and I know the theory of several supervised and unsupervised ML algorithms as well as deep learning. However, I lack huge amount of practical skills which is what I am working on right now. However, I am lost as to what to practice exactly. I am mostly interested in deep learning but I am seeing how essential it is to know how to implement other algorithms as well like random forests, SVM.. etc. Do I practice DL and other ML algorithms simultaneously (as mastering one area is almost impossible and is done across many years of experience) or should I first focus on one area (say computer vision) and then move on to the rest?
15 tech tips you won't find in a user manual
Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. Most gadgets don't come with a user manual that spells out every single feature. We learn them by doing, when someone spills the beans, or asking, "How'd you do that?" For example, no one thinks to dive into a new router's settings.
Deep Probabilistic Graphical Modeling
Probabilistic graphical modeling (PGM) provides a framework for formulating an interpretable generative process of data and expressing uncertainty about unknowns, but it lacks flexibility. Deep learning (DL) is an alternative framework for learning from data that has achieved great empirical success in recent years. DL offers great flexibility, but it lacks the interpretability and calibration of PGM. This thesis develops deep probabilistic graphical modeling (DPGM.) DPGM consists in leveraging DL to make PGM more flexible. DPGM brings about new methods for learning from data that exhibit the advantages of both PGM and DL. We use DL within PGM to build flexible models endowed with an interpretable latent structure. One model class we develop extends exponential family PCA using neural networks to improve predictive performance while enforcing the interpretability of the latent factors. Another model class we introduce enables accounting for long-term dependencies when modeling sequential data, which is a challenge when using purely DL or PGM approaches. Finally, DPGM successfully solves several outstanding problems of probabilistic topic models, a widely used family of models in PGM. DPGM also brings about new algorithms for learning with complex data. We develop reweighted expectation maximization, an algorithm that unifies several existing maximum likelihood-based algorithms for learning models parameterized by neural networks. This unifying view is made possible using expectation maximization, a canonical inference algorithm in PGM. We also develop entropy-regularized adversarial learning, a learning paradigm that deviates from the traditional maximum likelihood approach used in PGM. From the DL perspective, entropy-regularized adversarial learning provides a solution to the long-standing mode collapse problem of generative adversarial networks, a widely used DL approach.
Music Embedding: A Tool for Incorporating Music Theory into Computational Music Applications
HekmatiAthar, SeyyedPooya, Anwar, Mohd
Advancements in the digital technologies have enabled researchers to develop a variety of Computational Music applications. Such applications are required to capture, process, and generate data related to music. Therefore, it is important to digitally represent music in a music theoretic and concise manner. Existing approaches for representing music are ineffective in terms of utilizing music theory. In this paper, we address the disjoint of music theory and computational music by developing an opensource representation tool based on music theory. Through the wide range of use cases, we run an analysis on the classical music pieces to show the usefulness of the developed music embedding.
Kazuo Ishiguro writes of artificial intelligence and human hearts in 'Klara and the Sun'
Klara, the narrator of the new novel by Kazuo Ishiguro, isn't human, but understanding humans is her mission. In Klara and the Sun, the reader follows her in that mission, in a world that seems like our own in a none too distant future. Ishiguro, who was born in Japan but has lived most of his life in England, has written seven previous novels, including the Booker Prize-winning The Remains of the Day, as well as short fiction, song lyrics and screenplays. Klara and the Sun is his first novel since he received the Nobel Prize for literature in 2017. It underscores how well he deserved that prize, in its beautiful craft and prose and in its tender but unflinching sense of the human heart.