Media
ChatGPT, GPT-4, and More Generative AI News - KDnuggets
If you read my work you probably know that I publish my articles first and foremost in my AI newsletter, The Algorithmic Bridge. What you may not know is that every Sunday I publish a special column I call "what you may have missed," where I review everything that has happened during the week with analyses that help you make sense of the news. Semafor reported two weeks ago that, if everything goes according to the plan, Microsoft will close a $10B investment deal with OpenAI before the end of January (Satya Nadella, Microsoft's CEO, announced the extended partnership officially on Monday). There's been some misinformation about the deal which implied that OpenAI execs weren't sure about the company's long-term viability. Leo L'Orange, who writes The Neuron, explains that "once $92 billion in profit plus $13 billion in initial investment are repaid to Microsoft and once the other venture investors earn $150 billion, all of the equity reverts back to OpenAI."
The high-tech weeding machines cutting herbicide use
"AI is such a steep change in farming's evolution, it's like moving from ox to tractor," says Daniel McCann, who leads a company that is also infusing AI tech into a spraying solution for farmers. But with a twist in this case - his company, Precision AI, uses drones to fly over fields in the US Midwest to target weeds.
Practice of the conformer enhanced AUDIO-VISUAL HUBERT on Mandarin and English
Ren, Xiaoming, Li, Chao, Wang, Shenjian, Li, Biao
Considering the bimodal nature of human speech perception, lips, and teeth movement has a pivotal role in automatic speech recognition. Benefiting from the correlated and noise-invariant visual information, audio-visual recognition systems enhance robustness in multiple scenarios. In previous work, audio-visual HuBERT appears to be the finest practice incorporating modality knowledge. This paper outlines a mixed methodology, named conformer enhanced AV-HuBERT, boosting the AV-HuBERT system's performance a step further. Compared with baseline AV-HuBERT, our method in the one-phase evaluation of clean and noisy conditions achieves 7% and 16% relative WER reduction on the English AVSR benchmark dataset LRS3. Furthermore, we establish a novel 1000h Mandarin AVSR dataset CSTS. On top of the baseline AV-HuBERT, we exceed the WeNet ASR system by 14% and 18% relatively on MISP and CMLR by pre-training with this dataset. The conformer-enhanced AV-HuBERT we proposed brings 7% on MISP and 6% CER reduction on CMLR, compared with the baseline AV-HuBERT system.
Continuous descriptor-based control for deep audio synthesis
Devis, Ninon, Demerlรฉ, Nils, Nabi, Sarah, Genova, David, Esling, Philippe
Despite significant advances in deep models for music generation, the use of these techniques remains restricted to expert users. Before being democratized among musicians, generative models must first provide expressive control over the generation, as this conditions the integration of deep generative models in creative workflows. In this paper, we tackle this issue by introducing a deep generative audio model providing expressive and continuous descriptor-based control, while remaining lightweight enough to be embedded in a hardware synthesizer. We enforce the controllability of real-time generation by explicitly removing salient musical features in the latent space using an adversarial confusion criterion. User-specified features are then reintroduced as additional conditioning information, allowing for continuous control of the generation, akin to a synthesizer knob. We assess the performance of our method on a wide variety of sounds including instrumental, percussive and speech recordings while providing both timbre and attributes transfer, allowing new ways of generating sounds.
A Scalable Recommendation Engine for New Users and Items
Xu, Boya, Deng, Yiting, Mela, Carl
In many digital contexts such as online news and e-tailing with many new users and items, recommendation systems face several challenges: i) how to make initial recommendations to users with little or no response history (i.e., cold-start problem), ii) how to learn user preferences on items (test and learn), and iii) how to scale across many users and items with myriad demographics and attributes. While many recommendation systems accommodate aspects of these challenges, few if any address all. This paper introduces a Collaborative Filtering (CF) Multi-armed Bandit (B) with Attributes (A) recommendation system (CFB-A) to jointly accommodate all of these considerations. Empirical applications including an offline test on MovieLens data, synthetic data simulations, and an online grocery experiment indicate the CFB-A leads to substantial improvement on cumulative average rewards (e.g., total money or time spent, clicks, purchased quantities, average ratings, etc.) relative to the most powerful extant baseline methods.
A Comparative Analysis Of Latent Regressor Losses For Singing Voice Conversion
O'Connor, Brendan, Dixon, Simon
Previous research has shown that established techniques for spoken voice conversion (VC) do not perform as well when applied to singing voice conversion (SVC). We propose an alternative loss component in a loss function that is otherwise well-established among VC tasks, which has been shown to improve our model's SVC performance. We first trained a singer identity embedding (SIE) network on mel-spectrograms of singer recordings to produce singer-specific variance encodings using contrastive learning. We subsequently trained a well-known autoencoder framework (AutoVC) conditioned on these SIEs, and measured differences in SVC performance when using different latent regressor loss components. We found that using this loss w.r.t. SIEs leads to better performance than w.r.t. bottleneck embeddings, where converted audio is more natural and specific towards target singers. The inclusion of this loss component has the advantage of explicitly forcing the network to reconstruct with timbral similarity, and also negates the effect of poor disentanglement in AutoVC's bottleneck embeddings. We demonstrate peculiar diversity between computational and human evaluations on singer-converted audio clips, which highlights the necessity of both. We also propose a pitch-matching mechanism between source and target singers to ensure these evaluations are not influenced by differences in pitch register.
AI tool: The Future of Filmmaking. Generative A.I No Code
I've got some news for you, dear readers. It's time to prepare for the next big leap in filmmaking technology. And it comes in the form of an easy-to-use and free AI tool that will revolutionize the way we make films forever! This new tool promises to be an easy-to-use, free, and powerful platform that can automate your filmmaking process, thereby saving you valuable time and effort. We're talking about the latest in machine learning and artificial intelligence that's about to shake up the film industry as we know it. No more will we have to laboriously hand-craft every single visual effect or animation in our videos.
Five disturbing examples of why AI is not quite there
These AI mishaps show how far this new technology still has to go. The use of artificial intelligence is growing at a tremendous rate, especially with the recent release of OpenAI's chatbot ChatGPT. Although AI comes with its perks, it also comes with its mishaps. That has especially been proven true with OpenAI's other artificial intelligence invention known as DALL-E. CLICK TO GET KURT'S CYBERGUY NEWSLETTER WITH QUICK TIPS, TECH REVIEWS, SECURITY ALERTS AND EASY HOW-TO'S TO MAKE YOU SMARTER No, DALL-E is not the cousin of the beloved PIXAR robot WALL-E. DALL-E is a digital imaging learning model that was released back in 2021.
I'm Better Than Chatbots at the Job They're Trying to Take
If there is one thing the boosters and cynics agree on about artificial intelligence, it's that the tech is coming for white-collar jobs. This is not speculation of a far-off future--it's happening now. It makes sense, from a cold business perspective, that text-based media would want to adopt A.I. in order to cut costs (humans, expensive) and speed up output (humans, slow). Just look at how BuzzFeed's rock-bottom stock value jumped when it said last month that the site would use services from buzzy startup OpenAI to spiff up the site's famed quizzes. As Damon Beres wrote in the Atlantic shortly after the announcement: "The bleak future of media is human-owned websites profiting from automated banner ads placed on bot-written content, crawled by search-engine bots, and occasionally served to bot visitors."