Goto

Collaborating Authors

 Media


Modeling Perceptual Loudness of Piano Tone: Theory and Applications

arXiv.org Artificial Intelligence

The generation of piano tone The relationship between perceptual loudness and physical involves a complicated physical process [7] and the sound attributes of sound is an important subject in both computer contains rich timbral variations hard to be synthesized from music and psychoacoustics. Early studies of "equalloudness both frequency and time domain perspectives [8, 9]. On contour" can trace back to the 1920s and the measured the other hand, recently we see a growing number of music loudness with respect to intensity and frequency has information retrieval tasks involving feature extraction been revised many times since then. However, most studies of piano tone loudness, such as automatic music transcription merely focus on synthesized sound, and the induced [10-12] and performance rendering [13, 14]. In most theories on natural tones with complex timbre have rarely of these studies, loudness is sometimes confused with intensity been justified. To this end, we investigate both theory and or even the MIDI velocity. This motivates us to applications of natural-tone loudness perception in this paper investigate loudness perception specific to piano tone, beneficial via modeling piano tone. The theory part contains: for various downstream applications.


Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation

arXiv.org Artificial Intelligence

Symbolic music generation aims to generate music scores automatically. A recent trend is to use Transformer or its variants in music generation, which is, however, suboptimal, because the full attention cannot efficiently model the typically long music sequences (e.g., over 10,000 tokens), and the existing models have shortcomings in generating musical repetition structures. In this paper, we propose Museformer, a Transformer with a novel fine- and coarse-grained attention for music generation. Specifically, with the fine-grained attention, a token of a specific bar directly attends to all the tokens of the bars that are most relevant to music structures (e.g., the previous 1st, 2nd, 4th and 8th bars, selected via similarity statistics); with the coarse-grained attention, a token only attends to the summarization of the other bars rather than each token of them so as to reduce the computational cost. The advantages are two-fold. First, it can capture both music structure-related correlations via the fine-grained attention, and other contextual information via the coarse-grained attention. Second, it is efficient and can model over 3X longer music sequences compared to its full-attention counterpart. Both objective and subjective experimental results demonstrate its ability to generate long music sequences with high quality and better structures.


DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation

arXiv.org Artificial Intelligence

Dialog response generation in open domain is an important research topic where the main challenge is to generate relevant and diverse responses. In this paper, we propose a new dialog pre-training framework called DialogVED, which introduces continuous latent variables into the enhanced encoder-decoder pre-training framework to increase the relevance and diversity of responses. With the help of a large dialog corpus (Reddit), we pre-train the model using the following 4 tasks adopted in language models (LMs) and variational autoencoders (VAEs): 1) masked language model; 2) response generation; 3) bag-of-words prediction; and 4) KL divergence reduction. We also add additional parameters to model the turn structure in dialogs to improve the performance of the pre-trained model. We conduct experiments on PersonaChat, DailyDialog, and DSTC7-AVSD benchmarks for response generation. Experimental results show that our model achieves the new state-of-the-art results on all these datasets.


How artificial intelligence is changing music

#artificialintelligence

The final round of this year's AI Song Contest was contested by an oddly dissonant ode to coffee, some boppy Eurovision-esque ditties, a gently French melody, and a host of more genre-defying anthems. The competition is styled on the Eurovision, although it is open to entries from around the world, and the songs are entirely composed by computers. Listening to the finalists, and the ultimate winner โ€“ Thailand's Yaboi Hanoi with Asura Deva Choom Noom (Enter Demons & Gods); at aisongcontest.com โ€“ I found myself wondering if anyone had cheated by adding a little helping human hand. I also found myself wondering where the boundaries blur. Machines have been facilitating music-making since instruments were invented, but computer technology is a seismic shift on a par with the advents of sound recording and electrical amplification. While AI is currently causing future-shock rumblings, the influence of computers on our music has been felt for a while.


Image to text to music with CLIP interrogator and Mubert API

#artificialintelligence

The company Mubert is now venturing into a generative AI system that creates music based on text input. It is still in its infancy. Founded in 2017, U.S. startup Mubert specializes in generative AI for royalty-free music. Mubert's text-to-music app is a first attempt at generative AI that generates music from text input. A demo version at Huggingface allows users to input the prompt, from which the system then pulls individual keywords and matches them to the internal tagging of recorded sound clips, assembling a piece up to 100 seconds long.


Microsoft Open Sources Its 'Farm of the Future' Toolkit ยซ Machine Learning Times

#artificialintelligence

AI, artificial intelligence, data analytics, Deep Learning, Machine Learning, Predictive Analytics; 23 Views. Related. AI-Generated Imagery is the Newย โ€ฆ


Multiorder hydrologic Position for Europe -- a Set of Features for Machine Learning and โ€ฆ

#artificialintelligence

In the field of hydrogeology, machine learning has been used successfully for groundwater level prediction and a variety of mapping tasks3,4,5,6,7,8,9ย โ€ฆ


Citing 'New Needs,' Addiction Treatment Provider Aware Recovery Care Shakes Up C-Suite

#artificialintelligence

"Mark just brings 40 years of health information technology expertise, a lot of expertise in the area of machine learning and artificial โ€ฆ


Capgemini Acquires AI and Data Consultancy Quantmetry โ€“ ChannelE2E

#artificialintelligence

Quantmetry's areas of expertise include data mining, big data, machine learning, business intelligence, machine learning, computer vision, Python, โ€ฆ


Climate Nihilism--and Hope--Are Coming From the Strangest Places in Sci-Fi

Slate

Sign up to receive the Future Tense newsletter every other Saturday. The U.N.'s COP27 climate summit kicks off on Nov. 6 in Egypt, inviting us, once again, to consider whether we're doing enough, fast enough, to stave off climate chaos and the suffering that will come with it. The scale of change required is head-spinningly drastic, so even unexpectedly rapid expansions in clean energy won't do much to curb malaise and doomsaying. Here in the U.S., the Inflation Reduction Act, the biggest climate investment in the nation's history, has been met, largely, with collective indifference, despite positive buzz about its potential effectiveness. The bill was, predictably, passed without any Republican votes, a grim reminder of the scale of climate denialism.