Media
Multi-Task Learning Improves Performance In Deep Argument Mining Models
Farzam, Amirhossein, Shekhar, Shashank, Mehlhaff, Isaac, Morucci, Marco
The successful analysis of argumentative techniques from user-generated text is central to many downstream tasks such as political and market analysis. Recent argument mining tools use state-of-the-art deep learning methods to extract and annotate argumentative techniques from various online text corpora, however each task is treated as separate and different bespoke models are fine-tuned for each dataset. We show that different argument mining tasks share common semantic and logical structure by implementing a multi-task approach to argument mining that achieves better performance than state-of-the-art methods for the same problems. Our model builds a shared representation of the input text that is common to all tasks and exploits similarities between tasks in order to further boost performance via parameter-sharing. Our results are important for argument mining as they show that different tasks share substantial similarities and suggest a holistic approach to the extraction of argumentative techniques from text.
EmoGen: Eliminating Subjective Bias in Emotional Music Generation
Kang, Chenfei, Lu, Peiling, Yu, Botao, Tan, Xu, Ye, Wei, Zhang, Shikun, Bian, Jiang
Music is used to convey emotions, and thus generating emotional music is important in automatic music generation. Previous work on emotional music generation directly uses annotated emotion labels as control signals, which suffers from subjective bias: different people may annotate different emotions on the same music, and one person may feel different emotions under different situations. Therefore, directly mapping emotion labels to music sequences in an end-to-end way would confuse the learning process and hinder the model from generating music with general emotions. In this paper, we propose EmoGen, an emotional music generation system that leverages a set of emotion-related music attributes as the bridge between emotion and music, and divides the generation into two stages: emotion-to-attribute mapping with supervised clustering, and attribute-to-music generation with self-supervised learning. Both stages are beneficial: in the first stage, the attribute values around the clustering center represent the general emotions of these samples, which help eliminate the impacts of the subjective bias of emotion labels; in the second stage, the generation is completely disentangled from emotion labels and thus free from the subjective bias. Both subjective and objective evaluations show that EmoGen outperforms previous methods on emotion control accuracy and music quality respectively, which demonstrate our superiority in generating emotional music. Music samples generated by EmoGen are available via this link:https://ai-muzic.github.io/emogen/, and the code is available at this link:https://github.com/microsoft/muzic/.
Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video
Groh, Matthew, Sankaranarayanan, Aruna, Singh, Nikhil, Kim, Dong Young, Lippman, Andrew, Picard, Rosalind
Recent advances in technology for hyper-realistic visual effects provoke the concern that deepfake videos of political speeches will soon be visually indistinguishable from authentic video recordings. The conventional wisdom in communication theory predicts people will fall for fake news more often when the same version of a story is presented as a video versus text. We conduct 4 pre-registered randomized experiments with 2,015 participants to evaluate how accurately humans distinguish real political speeches from fabrications across base rates of misinformation, audio sources, and media modalities. We find base rates of misinformation minimally influence discernment and deepfakes with audio produced by the state-of-the-art text-to-speech algorithms are harder to discern than the same deepfakes with voice actor audio. Moreover, we find audio and visual information enables more accurate discernment than text alone: human discernment relies more on how something is said, the audio-visual cues, than what is said, the speech content.
Attention Mixtures for Time-Aware Sequential Recommendation
Tran, Viet-Anh, Salha-Galvan, Guillaume, Sguerra, Bruno, Hennequin, Romain
Transformers emerged as powerful methods for sequential recommendation. However, existing architectures often overlook the complex dependencies between user preferences and the temporal context. In this short paper, we introduce MOJITO, an improved Transformer sequential recommender system that addresses this limitation. MOJITO leverages Gaussian mixtures of attention-based temporal context and item embedding representations for sequential modeling. Such an approach permits to accurately predict which items should be recommended next to users depending on past actions and the temporal context. We demonstrate the relevance of our approach, by empirically outperforming existing Transformers for sequential recommendation on several real-world datasets.
An Overview on Language Models: Recent Developments and Outlook
Wei, Chengwei, Wang, Yun-Cheng, Wang, Bin, Kuo, C. -C. Jay
Language modeling studies the probability distributions over strings of texts. It is one of the most fundamental tasks in natural language processing (NLP). It has been widely used in text generation, speech recognition, machine translation, etc. Conventional language models (CLMs) aim to predict the probability of linguistic sequences in a causal manner, while pre-trained language models (PLMs) cover broader concepts and can be used in both causal sequential modeling and fine-tuning for downstream applications. PLMs have their own training paradigms (usually self-supervised) and serve as foundation models in modern NLP systems. This overview paper provides an introduction to both CLMs and PLMs from five aspects, i.e., linguistic units, architectures, training methods, evaluation methods, and applications. Furthermore, we discuss the relationship between CLMs and PLMs and shed light on the future directions of language modeling in the pre-trained era.
Valve won't publish games that feature copyright-infringing AI assets
Earlier this week, reports began surfacing that Valve was refusing to publish games with AI-generated art and other content. Over the weekend, the company finally commented on the matter. In a statement shared with IGN, Valve spokesperson Kaci Aitchison Boyle said the company is not trying to "discourage the use of [AI] on Steam." "Our priority, as always, is to try to ship as many of the titles we receive as we can," Aitchison Boyle said. "We welcome and encourage innovation, and AI technology is bound to create new and exciting experiences in gaming.
After "Barbie," Mattel Is Raiding Its Entire Toy Box
In 2019, Greta Gerwig became the latest in a line of writers, directors, and producers to make a pilgrimage to a toy workshop in El Segundo, California. Touring the facility, the Mattel Design Center, has become a rite of passage for Hollywood types who are considering transforming one of the company's products into a movie--a list that now includes such names as J. J. Abrams (Hot Wheels) and Vin Diesel (Rock'Em Sock'Em Robots). The building has hundreds of workspaces for artists, model-makers, and project managers, and it houses elaborate museum-style exhibitions that document the company's history and core products. These displays can help a toy designer find inspiration; they can also offer a "brand immersion"--a crash course in a Mattel property slated for adaptation. When a V.I.P. visits, Richard Dickson, a tall, bespectacled man who is the company's chief operating officer, plays the role of Willy Wonka. He'll show off the sixty-five-year-old machines that are still used to affix fake hair to Barbies; he'll invite you to inspect life-size, road-ready replicas of Hot Wheels cars. The center even boasts a giant rendering of Castle Grayskull, the fearsome ancestral home of He-Man.
Filter Bubbles in Recommender Systems: Fact or Fallacy -- A Systematic Review
Areeb, Qazi Mohammad, Nadeem, Mohammad, Sohail, Shahab Saquib, Imam, Raza, Doctor, Faiyaz, Himeur, Yassine, Hussain, Amir, Amira, Abbes
A filter bubble refers to the phenomenon where Internet customization effectively isolates individuals from diverse opinions or materials, resulting in their exposure to only a select set of content. This can lead to the reinforcement of existing attitudes, beliefs, or conditions. In this study, our primary focus is to investigate the impact of filter bubbles in recommender systems. This pioneering research aims to uncover the reasons behind this problem, explore potential solutions, and propose an integrated tool to help users avoid filter bubbles in recommender systems. To achieve this objective, we conduct a systematic literature review on the topic of filter bubbles in recommender systems. The reviewed articles are carefully analyzed and classified, providing valuable insights that inform the development of an integrated approach. Notably, our review reveals evidence of filter bubbles in recommendation systems, highlighting several biases that contribute to their existence. Moreover, we propose mechanisms to mitigate the impact of filter bubbles and demonstrate that incorporating diversity into recommendations can potentially help alleviate this issue. The findings of this timely review will serve as a benchmark for researchers working in interdisciplinary fields such as privacy, artificial intelligence ethics, and recommendation systems. Furthermore, it will open new avenues for future research in related domains, prompting further exploration and advancement in this critical area.
Large Language Models Enable Few-Shot Clustering
Viswanathan, Vijay, Gashteovski, Kiril, Lawrence, Carolin, Wu, Tongshuang, Neubig, Graham
Unlike traditional unsupervised clustering, semi-supervised clustering allows users to provide meaningful structure to the data, which helps the clustering algorithm to match the user's intent. Existing approaches to semi-supervised clustering require a significant amount of feedback from an expert to improve the clusters. In this paper, we ask whether a large language model can amplify an expert's guidance to enable query-efficient, few-shot semi-supervised text clustering. We show that LLMs are surprisingly effective at improving clustering. We explore three stages where LLMs can be incorporated into clustering: before clustering (improving input features), during clustering (by providing constraints to the clusterer), and after clustering (using LLMs post-correction). We find incorporating LLMs in the first two stages can routinely provide significant improvements in cluster quality, and that LLMs enable a user to make trade-offs between cost and accuracy to produce desired clusters. We release our code and LLM prompts for the public to use.