Media
MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies
Chen, Ke, Wu, Yusong, Liu, Haohe, Nezhurina, Marianna, Berg-Kirkpatrick, Taylor, Dubnov, Shlomo
Diffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited availability of music data and sensitive issues related to copyright and plagiarism. In this paper, to tackle these challenges, we first construct a state-of-the-art text-to-music model, MusicLDM, that adapts Stable Diffusion and AudioLDM architectures to the music domain. We achieve this by retraining the contrastive language-audio pretraining model (CLAP) and the Hifi-GAN vocoder, as components of MusicLDM, on a collection of music data samples. Then, to address the limitations of training data and to avoid plagiarism, we leverage a beat tracking model and propose two different mixup strategies for data augmentation: beat-synchronous audio mixup and beat-synchronous latent mixup, which recombine training audio directly or via a latent embeddings space, respectively. Such mixup strategies encourage the model to interpolate between musical training samples and generate new music within the convex hull of the training data, making the generated music more diverse while still staying faithful to the corresponding style. In addition to popular evaluation metrics, we design several new evaluation metrics based on CLAP score to demonstrate that our proposed MusicLDM and beat-synchronous mixup strategies improve both the quality and novelty of generated music, as well as the correspondence between input text and generated music.
Comparing scalable strategies for generating numerical perspectives
Cao, Hancheng, Spatharioti, Sofia Eleni, Goldstein, Daniel G., Hofman, Jake M.
Like other extreme quantities (for example, a distance of 34 parsecs), unfamiliar dollar amounts can be hard to fathom without comparison to something else [7, 18]. To address this issue, it can be useful to employ perspectives: re-phrasings of measurements that make them easier to understand, via a change of units to express the focal number on a different scale, or a comparison to a reference object. For instance, $330 billion can be re-expressed using perspectives of "about $1,000 per person in the United States" or "about 5% of the United States Federal Budget". In addition to being intuitively appealing and perceived as helpful [6, 10, 13], perspectives have been shown to aid numerical comprehension by boosting recall, estimation, error detection, and prediction [3, 12, 27], which could find relevance in a wide variety of downstream applications. These demonstrations of the benefits of perspectives have led to questions around what makes some analogies better than others, and if and how one can generate high-quality perspectives at scale for naturally occurring mentions of measurements. Approaches to automated perspective generation have varied, but they generally rely on first constructing a database of reference objects to compare measurements to and then prioritizing analogies to these reference objects that are both familiar and helpful to the reader [10, 27]. Prioritizing reference objects is complicated by the fact that what is most helpful for understanding a measurement can be difficult to quantify and can depend on the context in which the measurement occurs.
Joint Channel Estimation and Feedback with Masked Token Transformers in Massive MIMO Systems
Zhao, Mingming, Liu, Lin, Liu, Lifu, Li, Mengke, Tian, Qi
The downlink channel state information (CSI) estimation and low overhead acquisition are the major challenges for massive MIMO systems in frequency division duplex to enable high MIMO gain. Recently, numerous studies have been conducted to harness the power of deep neural networks for better channel estimation and feedback. However, existing methods have yet to fully exploit the intrinsic correlation features present in CSI. As a consequence, distinct network structures are utilized for handling these two tasks separately. To achieve joint channel estimation and feedback, this paper proposes an encoder-decoder based network that unveils the intrinsic frequency-domain correlation within the CSI matrix. The entire encoder-decoder network is utilized for channel compression. To effectively capture and restructure correlation features, a self-mask-attention coding is proposed, complemented by an active masking strategy designed to improve efficiency. The channel estimation is achieved through the decoder part, wherein a lightweight multilayer perceptron denoising module is utilized for further accurate estimation. Extensive experiments demonstrate that our method not only outperforms state-of-the-art channel estimation and feedback techniques in joint tasks but also achieves beneficial performance in individual tasks.
Mapping ChatGPT in Mainstream Media to Unravel Jobs and Diversity Challenges: Early Quantitative Insights through Sentiment Analysis and Word Frequency Analysis
The exponential growth in user acquisition and popularity of OpenAIs ChatGPT, an artificial intelligence(AI) powered chatbot, was accompanied by widespread mainstream media coverage. This article presents a quantitative data analysis of the early trends and sentiments revealed by conducting text mining and NLP methods onto a corpus of 10,902 mainstream news headlines related to the subject of ChatGPT and artificial intelligence, from the launch of ChatGPT in November 2022 to March 2023. The findings revealed in sentiment analysis, ChatGPT and artificial intelligence, were perceived more positively than negatively in the mainstream media. In regards to word frequency results, over sixty-five percent of the top frequency words were focused on Big Tech issues and actors while topics such as jobs, diversity, ethics, copyright, gender and women were poorly represented or completely absent and only accounted for six percent of the total corpus. This article is a critical analysis into the power structures and collusions between Big Tech and Big Media in their hegemonic exclusion of diversity and job challenges from mainstream media.
Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted Approach
Wu, Haoning, Zhang, Erli, Liao, Liang, Chen, Chaofeng, Hou, Jingwen, Wang, Annan, Sun, Wenxiu, Yan, Qiong, Lin, Weisi
The proliferation of in-the-wild videos has greatly expanded the Video Quality Assessment (VQA) problem. Unlike early definitions that usually focus on limited distortion types, VQA on in-the-wild videos is especially challenging as it could be affected by complicated factors, including various distortions and diverse contents. Though subjective studies have collected overall quality scores for these videos, how the abstract quality scores relate with specific factors is still obscure, hindering VQA methods from more concrete quality evaluations (e.g. sharpness of a video). To solve this problem, we collect over two million opinions on 4,543 in-the-wild videos on 13 dimensions of quality-related factors, including in-capture authentic distortions (e.g. motion blur, noise, flicker), errors introduced by compression and transmission, and higher-level experiences on semantic contents and aesthetic issues (e.g. composition, camera trajectory), to establish the multi-dimensional Maxwell database. Specifically, we ask the subjects to label among a positive, a negative, and a neutral choice for each dimension. These explanation-level opinions allow us to measure the relationships between specific quality factors and abstract subjective quality ratings, and to benchmark different categories of VQA algorithms on each dimension, so as to more comprehensively analyze their strengths and weaknesses. Furthermore, we propose the MaxVQA, a language-prompted VQA approach that modifies vision-language foundation model CLIP to better capture important quality issues as observed in our analyses. The MaxVQA can jointly evaluate various specific quality factors and final quality scores with state-of-the-art accuracy on all dimensions, and superb generalization ability on existing datasets. Code and data available at https://github.com/VQAssessment/MaxVQA.
Meaningful human command: Advance control directives as a method to enable moral and legal responsibility for autonomous weapons systems
21st Century war is increasing in speed, with conventional forces combined with massed use of autonomous systems and human-machine integration. However, a significant challenge is how humans can ensure moral and legal responsibility for systems operating outside of normal temporal parameters. This chapter considers whether humans can stand outside of real time and authorise actions for autonomous systems by the prior establishment of a contract, for actions to occur in a future context particularly in faster than real time or in very slow operations where human consciousness and concentration could not remain well informed. The medical legal precdent found in 'advance care directives' suggests how the time-consuming, deliberative process required for accountability and responsibility of weapons systems may be achievable outside real time captured in an 'advance control driective' (ACD). The chapter proposes 'autonomy command' scaffolded and legitimised through the construction of ACD ahead of the deployment of autonomous systems.
Matrix Estimation for Individual Fairness
Zhang, Cindy Y., Cen, Sarah H., Shah, Devavrat
In recent years, multiple notions of algorithmic fairness have arisen. One such notion is individual fairness (IF), which requires that individuals who are similar receive similar treatment. In parallel, matrix estimation (ME) has emerged as a natural paradigm for handling noisy data with missing values. In this work, we connect the two concepts. We show that pre-processing data using ME can improve an algorithm's IF without sacrificing performance. Specifically, we show that using a popular ME method known as singular value thresholding (SVT) to pre-process the data provides a strong IF guarantee under appropriate conditions. We then show that, under analogous conditions, SVT pre-processing also yields estimates that are consistent and approximately minimax optimal. As such, the ME pre-processing step does not, under the stated conditions, increase the prediction error of the base algorithm, i.e., does not impose a fairness-performance trade-off. We verify these results on synthetic and real data.
Sharing to learn and learning to share -- Fitting together Meta-Learning, Multi-Task Learning, and Transfer Learning: A meta review
Upadhyay, Richa, Phlypo, Ronald, Saini, Rajkumar, Liwicki, Marcus
Integrating knowledge across different domains is an essential feature of human learning. Learning paradigms such as transfer learning, meta learning, and multi-task learning reflect the human learning process by exploiting the prior knowledge for new tasks, encouraging faster learning and good generalization for new tasks. This article gives a detailed view of these learning paradigms and their comparative analysis. The weakness of one learning algorithm turns out to be a strength of another, and thus merging them is a prevalent trait in the literature. There are numerous research papers that focus on each of these learning paradigms separately and provide a comprehensive overview of them. However, this article provides a review of research studies that combine (two of) these learning algorithms. This survey describes how these techniques are combined to solve problems in many different fields of study, including computer vision, natural language processing, hyperspectral imaging, and many more, in supervised setting only. As a result, the global generic learning network an amalgamation of meta learning, transfer learning, and multi-task learning is introduced here, along with some open research questions and future research directions in the multi-task setting.
AI's impact on Hollywood amid the 'Barbenheimer' epic frenzy
CyberGuy shows how to take a sequence of action photos using Burst Mode and select the best one. So, you've probably heard about the latest stir in Hollywood โ and no, it's not about another celebrity feud or a blockbuster release. CLICK TO GET KURT'S FREE CYBERGUY NEWSLETTER WITH SECURITY ALERTS, QUICK TIPS, TECH REVIEWS AND EASY HOW-TO'S TO MAKE YOU SMARTER Instead, the buzz is all about artificial intelligence elbowing its way into the director's chair, churning out movie trailers, title sequences, and even entire episodes of beloved shows. It's an uncanny blend of technology and creativity that has everyone from industry insiders to casual moviegoers sitting up and taking notice. The "Barbie" movie's debut has been postponed in the Middle East until the end of August due to Warner Brothers still working on an edit of the film that will appease the region's censors.
Fruit-picking robots take flight, just when you've seen it all
Developed by the Israeli startup Tevel Aerobotics Technologies, these bots hover next to fruit trees, effortlessly pick the ripest fruits with suction arms, and carefully deposit them in a collection bin. With labor shortages leaving a wealth of fruit to rot on trees, farmers around the globe are seeing a glimmer of hope on the horizon as AI-driven robots fly to their rescue. These tech marvels are combating an escalating labor shortage that's leaving vast quantities of fruit to decay on trees. CLICK TO GET KURT'S FREE CYBERGUY NEWSLETTER WITH SECURITY ALERTS, QUICK TIPS, TECH REVIEWS AND EASY HOW-TO'S Developed by the Israeli startup Tevel Aerobotics Technologies, these bots hover next to fruit trees, effortlessly pick the ripest fruits with suction arms, and carefully deposit them in a collection bin. Like diligent honeybees, they're tethered to a platform that provides continuous power, enabling them to work day and night.