Media
Artificial intelligence news presenter Fedha started to work - Vita Gazette
Vita gazette โ Kuwait News virtual news server powered by artificial intelligence "Money" introduced. Artificial intelligence, which we encounter in almost every area, has also made rapid inroads into the media industry. It still needs to be determined whether he will make journalists unemployed in the future, but what is certain is that he is a rival. Appears as an image of a woman with light hair wearing a black jacket and white t-shirt'Money' Arabic said: "I'm Fedha, the first anchor to work with artificial intelligence at Kuwait News in Kuwait. Which news do you prefer? Kuwait News stated that Fedha was still in the testing phase and announced that it would read the daily news from the site's online bulletins. Abdullah Boftain, Deputy Editor-in-Chief of Kuwait News, 'Money'He said he would present the news with a Kuwaiti accent. About what'Money' He explained that they named it like this: "Fedha is a popular old Kuwaiti name, referring to silver and metal.
Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting
Wang, Su, Saharia, Chitwan, Montgomery, Ceslee, Pont-Tuset, Jordi, Noy, Shai, Pellegrini, Stefano, Onoe, Yasumasa, Laszlo, Sarah, Fleet, David J., Soricut, Radu, Baldridge, Jason, Norouzi, Mohammad, Anderson, Peter, Chan, William
Text-guided image editing can have a transformative impact in supporting creative applications. A key challenge is to generate edits that are faithful to input text prompts, while consistent with input images. We present Imagen Editor, a cascaded diffusion model built, by fine-tuning Imagen on text-guided image inpainting. Imagen Editor's edits are faithful to the text prompts, which is accomplished by using object detectors to propose inpainting masks during training. In addition, Imagen Editor captures fine details in the input image by conditioning the cascaded pipeline on the original high resolution image. To improve qualitative and quantitative evaluation, we introduce EditBench, a systematic benchmark for text-guided image inpainting. EditBench evaluates inpainting edits on natural and generated images exploring objects, attributes, and scenes. Through extensive human evaluation on EditBench, we find that object-masking during training leads to across-the-board improvements in text-image alignment -- such that Imagen Editor is preferred over DALL-E 2 and Stable Diffusion -- and, as a cohort, these models are better at object-rendering than text-rendering, and handle material/color/size attributes better than count/shape attributes.
Dynamic Mixed Membership Stochastic Block Model for Weighted Labeled Networks
Poux-Mรฉdard, Gaรซl, Velcin, Julien, Loudcher, Sabine
Most real-world networks evolve over time. Existing literature proposes models for dynamic networks that are either unlabeled or assumed to have a single membership structure. On the other hand, a new family of Mixed Membership Stochastic Block Models (MMSBM) allows to model static labeled networks under the assumption of mixed-membership clustering. In this work, we propose to extend this later class of models to infer dynamic labeled networks under a mixed membership assumption. Our approach takes the form of a temporal prior on the model's parameters. It relies on the single assumption that dynamics are not abrupt. We show that our method significantly differs from existing approaches, and allows to model more complex systems --dynamic labeled networks. We demonstrate the robustness of our method with several experiments on both synthetic and real-world datasets. A key interest of our approach is that it needs very few training data to yield good results. The performance gain under challenging conditions broadens the variety of possible applications of automated learning tools --as in social sciences, which comprise many fields where small datasets are a major obstacle to the introduction of machine learning methods.
Looking Similar, Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation Learning
Singh, Nikhil, Wu, Chih-Wei, Orife, Iroro, Kalayeh, Mahdi
Audiovisual representation learning typically relies on the correspondence between sight and sound. However, there are often multiple audio tracks that can correspond with a visual scene. Consider, for example, different conversations on the same crowded street. The effect of such counterfactual pairs on audiovisual representation learning has not been previously explored. To investigate this, we use dubbed versions of movies to augment cross-modal contrastive learning. Our approach learns to represent alternate audio tracks, differing only in speech content, similarly to the same video. Our results show that dub-augmented training improves performance on a range of auditory and audiovisual tasks, without significantly affecting linguistic task performance overall. We additionally compare this approach to a strong baseline where we remove speech before pretraining, and find that dub-augmented training is more effective, including for paralinguistic and audiovisual tasks where speech removal leads to worse performance. These findings highlight the importance of considering speech variation when learning scene-level audiovisual correspondences and suggest that dubbed audio can be a useful augmentation technique for training audiovisual models toward more robust performance.
A Phoneme-Informed Neural Network Model for Note-Level Singing Transcription
Yong, Sangeon, Su, Li, Nam, Juhan
Note-level automatic music transcription is one of the most representative music information retrieval (MIR) tasks and has been studied for various instruments to understand music. However, due to the lack of high-quality labeled data, transcription of many instruments is still a challenging task. In particular, in the case of singing, it is difficult to find accurate notes due to its expressiveness in pitch, timbre, and dynamics. In this paper, we propose a method of finding note onsets of singing voice more accurately by leveraging the linguistic characteristics of singing, which are not seen in other instruments. The proposed model uses mel-scaled spectrogram and phonetic posteriorgram (PPG), a frame-wise likelihood of phoneme, as an input of the onset detection network while PPG is generated by the pre-trained network with singing and speech data. To verify how linguistic features affect onset detection, we compare the evaluation results through the dataset with different languages and divide onset types for detailed analysis. Our approach substantially improves the performance of singing transcription and therefore emphasizes the importance of linguistic features in singing analysis.
On Distillation of Guided Diffusion Models
Meng, Chenlin, Rombach, Robin, Gao, Ruiqi, Kingma, Diederik P., Ermon, Stefano, Ho, Jonathan, Salimans, Tim
Classifier-free guided diffusion models have recently been shown to be highly effective at high-resolution image generation, and they have been widely used in large-scale diffusion frameworks including DALLE-2, Stable Diffusion and Imagen. However, a downside of classifier-free guided diffusion models is that they are computationally expensive at inference time since they require evaluating two diffusion models, a class-conditional model and an unconditional model, tens to hundreds of times. To deal with this limitation, we propose an approach to distilling classifier-free guided diffusion models into models that are fast to sample from: Given a pre-trained classifier-free guided model, we first learn a single model to match the output of the combined conditional and unconditional models, and then we progressively distill that model to a diffusion model that requires much fewer sampling steps. For standard diffusion models trained on the pixel-space, our approach is able to generate images visually comparable to that of the original model using as few as 4 sampling steps on ImageNet 64x64 and CIFAR-10, achieving FID/IS scores comparable to that of the original model while being up to 256 times faster to sample from. For diffusion models trained on the latent-space (e.g., Stable Diffusion), our approach is able to generate high-fidelity images using as few as 1 to 4 denoising steps, accelerating inference by at least 10-fold compared to existing methods on ImageNet 256x256 and LAION datasets. We further demonstrate the effectiveness of our approach on text-guided image editing and inpainting, where our distilled model is able to generate high-quality results using as few as 2-4 denoising steps.
A Scalable Framework for Automatic Playlist Continuation on Music Streaming Services
Bendada, Walid, Salha-Galvan, Guillaume, Bouabรงa, Thomas, Cazenave, Tristan
Music streaming services often aim to recommend songs for users to extend the playlists they have created on these services. However, extending playlists while preserving their musical characteristics and matching user preferences remains a challenging task, commonly referred to as Automatic Playlist Continuation (APC). Besides, while these services often need to select the best songs to recommend in real-time and among large catalogs with millions of candidates, recent research on APC mainly focused on models with few scalability guarantees and evaluated on relatively small datasets. In this paper, we introduce a general framework to build scalable yet effective APC models for large-scale applications. Based on a represent-then-aggregate strategy, it ensures scalability by design while remaining flexible enough to incorporate a wide range of representation learning and sequence modeling techniques, e.g., based on Transformers. We demonstrate the relevance of this framework through in-depth experimental validation on Spotify's Million Playlist Dataset (MPD), the largest public dataset for APC. We also describe how, in 2022, we successfully leveraged this framework to improve APC in production on Deezer. We report results from a large-scale online A/B test on this service, emphasizing the practical impact of our approach in such a real-world application.
China will require AI to reflect socialist values, not challenge social order
Fox News correspondent Matt Finn has the latest on the impact of AI technology that some say could outpace humans on'Special Report.' China on Tuesday revealed its proposed assessment measures for prospective generative artificial intelligence (AI) tools, telling companies they must submit their products before launching to the public. The Cyberspace Administration of China (CAC) proposed the measures in order to prevent discriminatory content, false information and content with the potential to harm personal privacy or intellectual property, the South China Morning Press reported. Such measures would ensure that the products do not end up suggesting regime subversion or disrupting economic or social order, according to the CAC. A number of Chinese companies, including Baidu, SenseTime and Alibaba, have recently shown of new AI models to power a number of applications from chatbots to image generators, prompting concern from officials over the impending boom in use.
Is em Beau Is Afraid /em Really a Comedy, or Is It As Scary As em Hereditary /em and em Midsommar /em ?
For die-hards, no horror movie can be too scary. But for you, a wimp, the wrong one can leave you miserable. Never fear, scaredies, because Slate's Scaredy Scale is here to help. We've put together a highly scientific and mostly spoiler-free system for rating new horror movies, comparing them with classics along a 10-point scale. And because not everyone is scared by the same things--some viewers can't stand jump scares, while others are haunted by more psychological terrors or can't stomach arterial spurts--it breaks down each movie's scares across three criteria: suspense, spookiness, and gore.
What Is AI Going To Do To Art? The History Of Photography Offers Clues.
Lois Rosson is a historian of science and technology based in Los Angeles. She is currently writing a book about images of outer space and their legibility. In 1835, William Henry Fox Talbot finally succeeded in producing a crude photograph of his country estate. He triumphantly declared that his was the first house ever known to have drawn its own picture. Fox Talbot described the calotype, his contribution to the photomechanical process, as an eradication of human intervention.