Goto

Collaborating Authors

 Media


Artificial Intelligence lending a helping hand to the virtual event industry

#artificialintelligence

Since the coronavirus pandemic took hold of our lives, virtual events have taken centre stage. With evolving time, virtual events are becoming the new normal, and AI is no longer something that business owners, marketers or organisers can afford to ignore. There is no denying that AI is proving to be one of the most important resources available for creating top-notch, successful events. But before we go any further, let's understand, what is AI? Most of us relate Artificial intelligence (AI) to sci-fi movies like Matrix or Inception.


NeuSaver: Neural Adaptive Power Consumption Optimization for Mobile Video Streaming

arXiv.org Artificial Intelligence

Video streaming services strive to support high-quality videos at higher resolutions and frame rates to improve the quality of experience (QoE). However, high-quality videos consume considerable amounts of energy on mobile devices. This paper proposes NeuSaver, which reduces the power consumption of mobile devices when streaming videos by applying an adaptive frame rate to each video chunk without compromising user experience. NeuSaver generates an optimal policy that determines the appropriate frame rate for each video chunk using reinforcement learning (RL). The RL model automatically learns the policy that maximizes the QoE goals based on previous observations. NeuSaver also uses an asynchronous advantage actor-critic algorithm to reinforce the RL model quickly and robustly. Streaming servers that support NeuSaver preprocesses videos into segments with various frame rates, which is similar to the process of creating videos with multiple bit rates in dynamic adaptive streaming over HTTP. NeuSaver utilizes the commonly used H.264 video codec. We evaluated NeuSaver in various experiments and a user study through four video categories along with the state-of-the-art model. Our experiments showed that NeuSaver effectively reduces the power consumption of mobile devices when streaming video by an average of 16.14% and up to 23.12% while achieving high QoE.


Alias-Free Generative Adversarial Networks

arXiv.org Artificial Intelligence

We observe that despite their hierarchical convolutional nature, the synthesis process of typical generative adversarial networks depends on absolute pixel coordinates in an unhealthy manner. This manifests itself as, e.g., detail appearing to be glued to image coordinates instead of the surfaces of depicted objects. We trace the root cause to careless signal processing that causes aliasing in the generator network. Interpreting all signals in the network as continuous, we derive generally applicable, small architectural changes that guarantee that unwanted information cannot leak into the hierarchical synthesis process. The resulting networks match the FID of StyleGAN2 but differ dramatically in their internal representations, and they are fully equivariant to translation and rotation even at subpixel scales. Our results pave the way for generative models better suited for video and animation.


Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

arXiv.org Artificial Intelligence

Large-scale vision and language representation learning has shown promising improvements on various vision-language tasks. Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens. Because the visual tokens and word tokens are unaligned, it is challenging for the multimodal encoder to learn image-text interactions. In this paper, we introduce a contrastive loss to ALign the image and text representations BEfore Fusing (ALBEF) them through cross-modal attention, which enables more grounded vision and language representation learning. Unlike most existing methods, our method does not require bounding box annotations nor high-resolution images. In order to improve learning from noisy web data, we propose momentum distillation, a self-training method which learns from pseudo-targets produced by a momentum model. We provide a theoretical analysis of ALBEF from a mutual information maximization perspective, showing that different training tasks can be interpreted as different ways to generate views for an image-text pair. ALBEF achieves state-of-the-art performance on multiple downstream vision-language tasks. On image-text retrieval, ALBEF outperforms methods that are pre-trained on orders of magnitude larger datasets. On VQA and NLVR$^2$, ALBEF achieves absolute improvements of 2.37% and 3.84% compared to the state-of-the-art, while enjoying faster inference speed. Code and pre-trained models are available at https://github.com/salesforce/ALBEF/.


MultiBench: Multiscale Benchmarks for Multimodal Representation Learning

arXiv.org Artificial Intelligence

Learning multimodal representations involves integrating information from multiple heterogeneous sources of data. It is a challenging yet crucial area with numerous real-world applications in multimedia, affective computing, robotics, finance, human-computer interaction, and healthcare. Unfortunately, multimodal research has seen limited resources to study (1) generalization across domains and modalities, (2) complexity during training and inference, and (3) robustness to noisy and missing modalities. In order to accelerate progress towards understudied modalities and tasks while ensuring real-world robustness, we release MultiBench, a systematic and unified large-scale benchmark spanning 15 datasets, 10 modalities, 20 prediction tasks, and 6 research areas. MultiBench provides an automated end-to-end machine learning pipeline that simplifies and standardizes data loading, experimental setup, and model evaluation. To enable holistic evaluation, MultiBench offers a comprehensive methodology to assess (1) generalization, (2) time and space complexity, and (3) modality robustness. MultiBench introduces impactful challenges for future research, including scalability to large-scale multimodal datasets and robustness to realistic imperfections. To accompany this benchmark, we also provide a standardized implementation of 20 core approaches in multimodal learning. Simply applying methods proposed in different research areas can improve the state-of-the-art performance on 9/15 datasets. Therefore, MultiBench presents a milestone in unifying disjoint efforts in multimodal research and paves the way towards a better understanding of the capabilities and limitations of multimodal models, all the while ensuring ease of use, accessibility, and reproducibility. MultiBench, our standardized code, and leaderboards are publicly available, will be regularly updated, and welcomes inputs from the community.


Scene-adaptive Knowledge Distillation for Sequential Recommendation via Differentiable Architecture Search

arXiv.org Artificial Intelligence

Sequential recommender systems (SRS) have become a research hotspot due to its power in modeling user dynamic interests and sequential behavioral patterns. To maximize model expressive ability, a default choice is to apply a larger and deeper network architecture, which, however, often brings high network latency when generating online recommendations. Naturally, we argue that compressing the heavy recommendation models into middle- or light- weight neural networks is of great importance for practical production systems. To realize such a goal, we propose AdaRec, a knowledge distillation (KD) framework which compresses knowledge of a teacher model into a student model adaptively according to its recommendation scene by using differentiable Neural Architecture Search (NAS). Specifically, we introduce a target-oriented distillation loss to guide the structure search process for finding the student network architecture, and a cost-sensitive loss as constraints for model size, which achieves a superior trade-off between recommendation effectiveness and efficiency. In addition, we leverage Earth Mover's Distance (EMD) to realize many-to-many layer mapping during knowledge distillation, which enables each intermediate student layer to learn from other intermediate teacher layers adaptively. Extensive experiments on real-world recommendation datasets demonstrate that our model achieves competitive or better accuracy with notable inference speedup comparing to strong counterparts, while discovering diverse neural architectures for sequential recommender models under different recommendation scenes.


How AI is Transforming Music Composition -- Xyonix, AI Consulting & Custom Solutions

#artificialintelligence

Music is central to modern entertainment and culture. It serves as stand-alone art, as a backdrop to daily life, and as an important feature in industries like film and advertising. What if the music you hear on the radio was not composed by a human at all, but by a machine? The application of AI in creative spaces like music and art is not new, but recent years have seen a dramatic expanse in machine learning capabilities in music composition. AI is being used by researchers and startups to compose soundtracks and soundscapes, and to create original songs within the style of specific genres and artists.


Americans are turning to dating apps to find friends

FOX News

Fox News Flash top entertainment and celebrity headlines are here. Check out what's clicking today in entertainment. You might consider trying a dating app. That's what 26-year-old Gaby Deimeke did after she moved to Austin, Texas, in 2019. After hearing about Bumble BFF at a music festival, Deimeke download the app and gave it a try.


How AI can be used to streamline DAM workflows

#artificialintelligence

It wasn't too long ago that artificial intelligence (AI) technology was merely a concept from science fiction movies. But today, AI is changing the status quo of entire industries, including marketing technology (martech). In the marketing sphere, we are increasingly seeing AI capabilities leveraged in martech tools: the use of automation capabilities used to support chatbots, personalize content, and even manage social media platforms. As consumer expectations shift, and agile, disruptive retail players enter the ecosystem, marketers need to ensure that they are still able to compete for consumers' attention across an increasing number of channels. For these marketing teams, this means ensuring any visual assets used are both accurate and compelling enough in a competitive landscape to convert consumers into paying customers.


The Rise Of The Machines; Analogue Meets Artificial intelligence

#artificialintelligence

Sonnant styles itself as a "transformational artificial intelligence (AI) and machine learning (ML) company that provides content discovery for the …