Media
An investigation of the reconstruction capacity of stacked convolutional autoencoders for log-mel-spectrograms
Natsiou, Anastasia, Longo, Luca, O'Leary, Sean
In audio processing applications, the generation of expressive sounds based on high-level representations demonstrates a high demand. These representations can be used to manipulate the timbre and influence the synthesis of creative instrumental notes. Modern algorithms, such as neural networks, have inspired the development of expressive synthesizers based on musical instrument timbre compression. Unsupervised deep learning methods can achieve audio compression by training the network to learn a mapping from waveforms or spectrograms to low-dimensional representations. This study investigates the use of stacked convolutional autoencoders for the compression of time-frequency audio representations for a variety of instruments for a single pitch. Further exploration of hyper-parameters and regularization techniques is demonstrated to enhance the performance of the initial design. In an unsupervised manner, the network is able to reconstruct a monophonic and harmonic sound based on latent representations. In addition, we introduce an evaluation metric to measure the similarity between the original and reconstructed samples. Evaluating a deep generative model for the synthesis of sound is a challenging task. Our approach is based on the accuracy of the generated frequencies as it presents a significant metric for the perception of harmonic sounds. This work is expected to accelerate future experiments on audio compression using neural autoencoders.
InstructPix2Pix: Learning to Follow Image Editing Instructions
Brooks, Tim, Holynski, Aleksander, Efros, Alexei A.
We propose a method for editing images from human instructions: given an input image and a written instruction that tells the model what to do, our model follows these instructions to edit the image. To obtain training data for this problem, we combine the knowledge of two large pretrained models -- a language model (GPT-3) and a text-to-image model (Stable Diffusion) -- to generate a large dataset of image editing examples. Our conditional diffusion model, InstructPix2Pix, is trained on our generated data, and generalizes to real images and user-written instructions at inference time. Since it performs edits in the forward pass and does not require per example fine-tuning or inversion, our model edits images quickly, in a matter of seconds. We show compelling editing results for a diverse collection of input images and written instructions.
From Taylor Swift to David Bowie and Elvis Presley: AI technology creates new songs for musicians
The artificial intelligence bot which Nick Cave accused of making a'grotesque mockery' of his work has hit back at the musician by insisting it'tries its best to generate text that is coherent, creative and conveys a message'. Cave left a scathing review of ChatGPT's rendition of his work, describing the lyrics after fans asked it to replicate his style of music. He's one of many musicians fans are asking the bot to mimic, to see if it can capture the magic of songs legitimately produced by the artist. ChatGPT collates huge swathes of data which allows it to predict phrases and words that are likely to be used in an artist's repertoire. When provided a prompt, such as asking it to create song lyrics for a certain artist, it sweeps the database for all known previous works by the artist and collates sentences using phrases, terms and themes frequently used in association with the musician.
This 22-year-old is trying to save us from ChatGPT before it changes writing forever
While many Americans were nursing hangovers on New Year's Day, 22-year-old Edward Tian was working feverishly on a new app to combat misuse of a powerful, new artificial intelligence tool called ChatGPT. Given the buzz it's created, there's a good chance you've heard about ChatGPT. It's an interactive chatbot powered by machine learning. The technology has basically devoured the entire Internet, reading the collective works of humanity and learning patterns in language that it can recreate. All you have to do is give it a prompt, and ChatGPT can do an endless array of things: write a story in a particular style, answer a question, explain a concept, compose an email -- write a college essay -- and it will spit out coherent, seemingly human-written text in seconds.
Everyday AI podcast series
In a new podcast series, Everyday AI, host Jon Whittle (CSIRO) explores the AI that is already shaping our lives. With the help of expert guests, he explores how AI is used in creative industries, health, conservation, sports and space. Episode 4: AI and citizen science – AI in ecology This episode features Jessie Barry from Cornell University's Macaulay Library and Merlin Bird ID, ichthyologist Mark McGrouther, and Google's Megha Malpani. Episode 6: The final frontier – AI in space This episode features Astrophysicist Kirsten Banks, NASA researcher Dr Raymond Francis, and Research Astronomer Dr Ivy Wong.
Hytera Enhances New Generation H-Series DMR Two-way Radio with HP5 Models
Hytera Communications a leading global provider of professional communications technologies and solutions,released HP56X and HP50X portable two-way radios to further expand and strengthen its new generation of Digital Mobile Radio (DMR) portfolio. The HP5 models are developed to provide reliable voice communications for security, operations, technician, and maintenance teams at office buildings, stadiums, industrial parks, school campuses, hospitals, etc. "However, their requirements for versatility, ergonomics, and reliability are similar. With this in mind, we designed HP5 portable radios. We believe HP5 will be a great productivity and safety tool for a lot of professional scenarios." H-Series, including portable radios, mobile radios, and repeaters, is designed and developed on new hardware and software platforms.
How "Battle Royale" Took Over Video Games
In the mid-nineteen-nineties, Koushun Takami was dozing on his futon on the island of Shikoku, Japan, when he was visited by an apparition: a maniacal schoolteacher addressing a group of students. "All right, class, listen up," Takami heard the teacher say. "Today, I'm going to have you all kill each other." Takami was in his twenties, and he had recently quit his job as a reporter for a local newspaper to become a novelist. As a literature student at Osaka University, he had started and abandoned several horror-infused detective stories.
Top Image Editing Tools In 2022 - Fronty
A well-designed image can help you get more clicks and rank higher in search engines. However, the image should have a professional appearance. The image editing tools are helpful and make it simple to bring your image to life. Image Editing Tools allow you to create images with various creative effects to choose from, whether you're editing images for your website, social media campaigns, or other purposes. You can use these image editing tools to change the colors of images, change the background, and update text data.
AI Chatbot Writes 'In the Style of Nick Cave,' and Nick Cave is Heated – Rolling Stone
Nick Cave, the Bad Seeds frontman whose songs are tinged with a healthy dose of death, forlorn love, and religion, is no fan of ChatGPT's lyrical ambitions. The popular AI bot has drawn both praise and concern for its ability to generate conversational and nuanced text responses in simple, clean sentences. Since its release in November by the artificial intelligence lab OpenAI, ChatGPT has written everything from sitcom scripts to literature essays to, now, rather convincing rock songs. This has left people worried about the ramifications for industries across the creative spectrum, and one of those people is Cave himself. In his latest The Red Hand Files newsletter, Cave took on the subject of AI generated music.