Goto

Collaborating Authors

 Media


Meta's open-source MusicGen AI uses text to create song genre mashups

Engadget

Meta's Audiocraft research team has just released MusicGen, an open source deep learning language model that can generate new music based on text prompts and even be aligned to an existing song, The Decoder reported. It's much like ChatGPT for audio, letting you describe the style of music you want, drop in an existing tune (optionally) and then clicking "Generate." After a good chunk of time (around 160 seconds in my case), it spits out a short piece of all-new music based on your text prompts and melody. The demo on Facebook's Hugging Face AI site lets you describe your music, providing a handful of examples like "an 80s driving pop song with heavy drums and synth pads in the background." You can then "condition" that on a given song up top 30 seconds long, with controls letting select a specific portion of that.


When I lost my job, I learned to code. Now AI doom mongers are trying to scare me all over again Tristan Cross

The Guardian

I spent the best part of the 2010s working in new media, which โ€“ if you enjoyed being repeatedly laid off and then being inundated with jeering messages inveigling you to "learn to code" because your industry was doomed โ€“ was a great big laugh. Eventually, the fun began to wear off and in an act of subversive defiance (or cowardly resignation), I took their goading advice, learned to code and pivoted to what I'd hoped would be a far more secure career in "web development", only for recent advances in AI to supposedly render coding jobs a waste of time, too. It seems I have accidentally timed my career change to coincide with a mass rollout of AI chatbots that have also learned to code, and that are โ€“ in many respects โ€“ already far better at it than me. Code can appear alarming to the uninitiated: inscrutable "languages" that mostly read like a calculator having a stroke, but, according to AI's most fervent evangelists, they no longer need represent any barrier at all. Why bother wrapping your head around the needlessly convoluted nerdspeak required to display white text on a black background, when you can now simply ask a chatbot to do this in layperson's terms and it will promptly serve up your code, complete with instructions?


AI 'deepfakes' of innocent images fuel spike in sextortion scams, FBI warns

FOX News

Brian Montgomery, who lost his son to suicide after he was extorted, discussed the loss of his son and how teen boys have been blackmailed over explicit pictures on'America's Newsroom.' If you or someone you know is having thoughts of suicide, please contact the Suicide & Crisis Lifeline at 988 or 1-800-273-TALK (8255). Artifical intelligence-generated "deepfakes" are fueling sextortion scams like dried up brush in an out-of-control wildfire. The number of nationally reported sextortion cases increased 322% between February 2022 and February 2023, according to the FBI, which said last week there's been a significant uptick since April because of AI-doctored images. Innocent pictures or videos uploaded to social media or sent in messages can be twisted into sexually explicit, AI-generated images that are "true-to-life" and nearly impossible to discern, the FBI said.


Lost in Translation: Large Language Models in Non-English Content Analysis

arXiv.org Artificial Intelligence

In recent years, large language models (e.g., Open AI's GPT-4, Meta's LLaMa, Google's PaLM) have become the dominant approach for building AI systems to analyze and generate language online. However, the automated systems that increasingly mediate our interactions online -- such as chatbots, content moderation systems, and search engines -- are primarily designed for and work far more effectively in English than in the world's other 7,000 languages. Recently, researchers and technology companies have attempted to extend the capabilities of large language models into languages other than English by building what are called multilingual language models. In this paper, we explain how these multilingual language models work and explore their capabilities and limits. Part I provides a simple technical explanation of how large language models work, why there is a gap in available data between English and other languages, and how multilingual language models attempt to bridge that gap. Part II accounts for the challenges of doing content analysis with large language models in general and multilingual language models in particular. Part III offers recommendations for companies, researchers, and policymakers to keep in mind when considering researching, developing and deploying large and multilingual language models.


MSSRNet: Manipulating Sequential Style Representation for Unsupervised Text Style Transfer

arXiv.org Artificial Intelligence

Unsupervised text style transfer task aims to rewrite a text into target style while preserving its main content. Traditional methods rely on the use of a fixed-sized vector to regulate text style, which is difficult to accurately convey the style strength for each individual token. In fact, each token of a text contains different style intensity and makes different contribution to the overall style. Our proposed method addresses this issue by assigning individual style vector to each token in a text, allowing for fine-grained control and manipulation of the style strength. Additionally, an adversarial training framework integrated with teacher-student learning is introduced to enhance training stability and reduce the complexity of high-dimensional optimization. The results of our experiments demonstrate the efficacy of our method in terms of clearly improved style transfer accuracy and content preservation in both two-style transfer and multi-style transfer settings.


Towards Fair and Explainable AI using a Human-Centered AI Approach

arXiv.org Artificial Intelligence

The rise of machine learning (ML) is accompanied by several high-profile cases that have stressed the need for fairness, accountability, explainability and trust in ML systems. The existing literature has largely focused on fully automated ML approaches that try to optimize for some performance metric. However, human-centric measures like fairness, trust, explainability, etc. are subjective in nature, context-dependent, and might not correlate with conventional performance metrics. To deal with these challenges, we explore a human-centered AI approach that empowers people by providing more transparency and human control. In this dissertation, we present 5 research projects that aim to enhance explainability and fairness in classification systems and word embeddings. The first project explores the utility/downsides of introducing local model explanations as interfaces for machine teachers (crowd workers). Our study found that adding explanations supports trust calibration for the resulting ML model and enables rich forms of teaching feedback. The second project presents D-BIAS, a causality-based human-in-the-loop visual tool for identifying and mitigating social biases in tabular datasets. Apart from fairness, we found that our tool also enhances trust and accountability. The third project presents WordBias, a visual interactive tool that helps audit pre-trained static word embeddings for biases against groups, such as females, or subgroups, such as Black Muslim females. The fourth project presents DramatVis Personae, a visual analytics tool that helps identify social biases in creative writing. Finally, the last project presents an empirical study aimed at understanding the cumulative impact of multiple fairness-enhancing interventions at different stages of the ML pipeline on fairness, utility and different population groups. We conclude by discussing some of the future directions.


Fill-Up: Balancing Long-Tailed Data with Generative Models

arXiv.org Artificial Intelligence

Modern text-to-image synthesis models have achieved an exceptional level of photorealism, generating high-quality images from arbitrary text descriptions. In light of the impressive synthesis ability, several studies have exhibited promising results in exploiting generated data for image recognition. However, directly supplementing data-hungry situations in the real-world (e.g. few-shot or long-tailed scenarios) with existing approaches result in marginal performance gains, as they suffer to thoroughly reflect the distribution of the real data. Through extensive experiments, this paper proposes a new image synthesis pipeline for long-tailed situations using Textual Inversion. The study demonstrates that generated images from textual-inverted text tokens effectively aligns with the real domain, significantly enhancing the recognition ability of a standard ResNet50 backbone. We also show that real-world data imbalance scenarios can be successfully mitigated by filling up the imbalanced data with synthetic images. In conjunction with techniques in the area of long-tailed recognition, our method achieves state-of-the-art results on standard long-tailed benchmarks when trained from scratch.


Video-to-Music Recommendation using Temporal Alignment of Segments

arXiv.org Artificial Intelligence

We study cross-modal recommendation of music tracks to be used as soundtracks for videos. This problem is known as the music supervision task. We build on a self-supervised system that learns a content association between music and video. In addition to the adequacy of content, adequacy of structure is crucial in music supervision to obtain relevant recommendations. We propose a novel approach to significantly improve the system's performance using structure-aware recommendation. The core idea is to consider not only the full audio-video clips, but rather shorter segments for training and inference. We find that using semantic segments and ranking the tracks according to sequence alignment costs significantly improves the results. We investigate the impact of different ranking metrics and segmentation methods.


Evaluating the Social Impact of Generative AI Systems in Systems and Society

arXiv.org Artificial Intelligence

Generative AI systems across modalities, ranging from text, image, audio, and video, have broad social impacts, but there exists no official standard for means of evaluating those impacts and which impacts should be evaluated. We move toward a standard approach in evaluating a generative AI system for any modality, in two overarching categories: what is able to be evaluated in a base system that has no predetermined application and what is able to be evaluated in society. We describe specific social impact categories and how to approach and conduct evaluations in the base technical system, then in people and society. Our framework for a base system defines seven categories of social impact: bias, stereotypes, and representational harms; cultural values and sensitive content; disparate performance; privacy and data protection; financial costs; environmental costs; and data and content moderation labor costs. Suggested methods for evaluation apply to all modalities and analyses of the limitations of existing evaluations serve as a starting point for necessary investment in future evaluations. We offer five overarching categories for what is able to be evaluated in society, each with their own subcategories: trustworthiness and autonomy; inequality, marginalization, and violence; concentration of authority; labor and creativity; and ecosystem and environment. Each subcategory includes recommendations for mitigating harm. We are concurrently crafting an evaluation repository for the AI research community to contribute existing evaluations along the given categories. This version will be updated following a CRAFT session at ACM FAccT 2023.


Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

arXiv.org Artificial Intelligence

Artificial agents have traditionally been trained to maximize reward, which may incentivize power-seeking and deception, analogous to how next-token prediction in language models (LMs) may incentivize toxicity. So do agents naturally learn to be Machiavellian? And how do we measure these behaviors in general-purpose models such as GPT-4? Towards answering these questions, we introduce MACHIAVELLI, a benchmark of 134 Choose-Your-Own-Adventure games containing over half a million rich, diverse scenarios that center on social decision-making. Scenario labeling is automated with LMs, which are more performant than human annotators. We mathematize dozens of harmful behaviors and use our annotations to evaluate agents' tendencies to be power-seeking, cause disutility, and commit ethical violations. We observe some tension between maximizing reward and behaving ethically. To improve this trade-off, we investigate LM-based methods to steer agents' towards less harmful behaviors. Our results show that agents can both act competently and morally, so concrete progress can currently be made in machine ethics--designing agents that are Pareto improvements in both safety and capabilities.