Media
Out-of-Distribution Detection using Synthetic Data Generation
Abbas, Momin, Azmat, Muneeza, Horesh, Raya, Yurochkin, Mikhail
Distinguishing in- and out-of-distribution (OOD) inputs is crucial for reliable deployment of classification systems. However, OOD data is typically unavailable or difficult to collect, posing a significant challenge for accurate OOD detection. In this work, we present a method that harnesses the generative capabilities of Large Language Models (LLMs) to create high-quality synthetic OOD proxies, eliminating the dependency on any external OOD data source. We study the efficacy of our method on classical text classification tasks such as toxicity detection and sentiment classification as well as classification tasks arising in LLM development and deployment, such as training a reward model for RLHF and detecting misaligned generations. Extensive experiments on nine InD-OOD dataset pairs and various model sizes show that our approach dramatically lowers false positive rates (achieving a perfect zero in some cases) while maintaining high accuracy on in-distribution tasks, outperforming baseline methods by a significant margin.
Lowering the Barrier of Machine Learning: Achieving Zero Manual Labeling in Review Classification Using LLMs
With the internet's evolution, consumers increasingly rely on online reviews for service or product choices, necessitating that businesses analyze extensive customer feedback to enhance their offerings. While machine learning-based sentiment classification shows promise in this realm, its technical complexity often bars small businesses and individuals from leveraging such advancements, which may end up making the competitive gap between small and large businesses even bigger in terms of improving customer satisfaction. This paper introduces an approach that integrates large language models (LLMs), specifically Generative Pre-trained Transformer (GPT) and Bidirectional Encoder Representations from Transformers (BERT)-based models, making it accessible to a wider audience. Our experiments across various datasets confirm that our approach retains high classification accuracy without the need for manual labeling, expert knowledge in tuning and data annotation, or substantial computational power. By significantly lowering the barriers to applying sentiment classification techniques, our methodology enhances competitiveness and paves the way for making machine learning technology accessible to a broader audience.
Recommendations Beyond Catalogs: Diffusion Models for Personalized Generation
Patron, Gabriel, Xu, Zhiwei, Kapnadak, Ishan, Polo, Felipe Maia
Modern recommender systems follow the guiding principle of serving the right user, the right item at the right time. One of their main limitations is that they are typically limited to items already in the catalog. We propose REcommendations BEyond CAtalogs, REBECA, a new class of probabilistic diffusion-based recommender systems that synthesize new items tailored to individual tastes rather than retrieve items from the catalog. REBECA combines efficient training in embedding space with a novel diffusion prior that only requires users' past ratings of items. We evaluate REBECA on real-world data and propose novel personalization metrics for generative recommender systems. Extensive experiments demonstrate that REBECA produces high-quality, personalized recommendations, generating images that align with users' unique preferences.
This Clever New Book About the Apocalypse Will Cheer You Up (Really!)
So long as we can say'This is the worst,' " go the lines from King Lear quoted in Emily St. John Mandel's 2014 novel Station Eleven. Any stories we tell about the end of the world will have to be fictional, since once the real thing occurs, no one will be around to describe it. As the British journalist Dorian Lynskey relates in his erudite, delightfully witty, and strangely cheering new book, Everything Must Go: The Stories We Tell About the End of the World, the fact that we can only ever speculate on the subject makes us speculate all the more frantically. "There is simply no end of ends," Lynskey writes of the books, movies, TV shows, pop songs, and video games we've created to depict the apocalypse--or its near misses and the aftermaths thereof. Station Eleven is often described as "postapocalyptic," but as Lynskey points out, the more accurate term would be "postcatastrophic." That's a better label for stories in which "the world has not ended, but a world has, creating a blank ...
Does AI need all that money? (Tech giants say yes)
It's been another wild few days in Elon Musk news. Stay tuned for our coverage. In personal news, I deleted Instagram from my phone to try out a month without it there. Instead of scrolling, I've been listening to Shygirl and Lady Gaga's new music. DeepSeek roiled the US stock market last week by proposing that AI shouldn't really be all that expensive. The suggestion was so stunning it wiped about 600bn off of Nvidia's market cap in one day.
MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music Composition
Pasquier, Philippe, Ens, Jeff, Fradet, Nathan, Triana, Paul, Rizzotti, Davide, Rolland, Jean-Baptiste, Safi, Maryam
We present and release MIDI-GPT, a generative system based on the Transformer architecture that is designed for computer-assisted music composition workflows. MIDI-GPT supports the infilling of musical material at the track and bar level, and can condition generation on attributes including: instrument type, musical style, note density, polyphony level, and note duration. In order to integrate these features, we employ an alternative representation for musical material, creating a time-ordered sequence of musical events for each track and concatenating several tracks into a single sequence, rather than using a single time-ordered sequence where the musical events corresponding to different tracks are interleaved. We also propose a variation of our representation allowing for expressiveness. We present experimental results that demonstrate that MIDI-GPT is able to consistently avoid duplicating the musical material it was trained on, generate music that is stylistically similar to the training dataset, and that attribute controls allow enforcing various constraints on the generated material. We also outline several real-world applications of MIDI-GPT, including collaborations with industry partners that explore the integration and evaluation of MIDI-GPT into commercial products, as well as several artistic works produced using it.
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
Liang, Feng, Ma, Haoyu, He, Zecheng, Hou, Tingbo, Hou, Ji, Li, Kunpeng, Dai, Xiaoliang, Juefei-Xu, Felix, Azadi, Samaneh, Sinha, Animesh, Zhang, Peizhao, Vajda, Peter, Marculescu, Diana
Video personalization, which generates customized videos using reference images, has gained significant attention. However, prior methods typically focus on single-concept personalization, limiting broader applications that require multi-concept integration. Attempts to extend these models to multiple concepts often lead to identity blending, which results in composite characters with fused attributes from multiple sources. This challenge arises due to the lack of a mechanism to link each concept with its specific reference image. We address this with anchored prompts, which embed image anchors as unique tokens within text prompts, guiding accurate referencing during generation. Additionally, we introduce concept embeddings to encode the order of reference images. Our approach, Movie Weaver, seamlessly weaves multiple concepts-including face, body, and animal images-into one video, allowing flexible combinations in a single model. The evaluation shows that Movie Weaver outperforms existing methods for multi-concept video personalization in identity preservation and overall quality.
Mass-Editing Memory with Attention in Transformers: A cross-lingual exploration of knowledge
Tamayo, Daniel, Gonzalez-Agirre, Aitor, Hernando, Javier, Villegas, Marta
Recent research has explored methods for updating and modifying factual knowledge in large language models, often focusing on specific multi-layer perceptron blocks. This study expands on this work by examining the effectiveness of existing knowledge editing methods across languages and delving into the role of attention mechanisms in this process. Drawing from the insights gained, we propose Mass-Editing Memory with Attention in Transformers (MEMAT), a method that achieves significant improvements in all metrics while requiring minimal parameter modifications. MEMAT delivers a remarkable 10% increase in magnitude metrics, benefits languages not included in the training data and also demonstrates a high degree of portability. Our code and data are at https://github.com/dtamayo-nlp/MEMAT.
Embracing Dialectic Intersubjectivity: Coordination of Different Perspectives in Content Analysis with LLM Persona Simulation
Kang, Taewoo, Thorson, Kjerstin, Peng, Tai-Quan, Hiaeshutter-Rice, Dan, Lee, Sanguk, Soroka, Stuart
Large Language Models (LLMs), such as ChatGPT, allow researchers to analyze extensive text corpora with efficiency (Bail, 2024). However, human and algorithmic biases can influence their outputs, reflecting ideological or cultural skewness in the training data (Bender et al., 2021; Kroon et al., 2023). With the growing adoption of LLM-Assisted Content Analysis (LACA)--a method replacing manual coding with LLM-generated datasets (Chew et al., 2023)--the scientific community must ensure these tools enhance understanding rather than reinforce existing biases (Messeri and Crockett, 2024). Human bias is a long-standing issue in content analysis, which traditionally depends on a consensus-oriented research practice. Intercoder agreement is used to ensure reliability, aiming for shared interpretations of text that can be statistically validated (Krippendorff, 1999). However, this emphasis on consensus may inadvertently neglect diverse perspectives, favoring uniform interpretations that risk simplifying the sociocultural complexities embedded in textual connotations.
ReSpark: Leveraging Previous Data Reports as References to Generate New Reports with LLMs
Tian, Yuan, Zhang, Chuhan, Wang, Xiaotong, Pan, Sitong, Cui, Weiwei, Zhang, Haidong, Deng, Dazhen, Wu, Yingcai
Creating data reports is time-consuming, as it requires iterative exploration and understanding of data, followed by summarizing the insights. While large language models (LLMs) are powerful tools for data processing and text generation, they often struggle to produce complete data reports that fully meet user expectations. One significant challenge is effectively communicating the entire analysis logic to LLMs. Moreover, determining a comprehensive analysis logic can be mentally taxing for users. To address these challenges, we propose ReSpark, an LLM-based method that leverages existing data reports as references for creating new ones. Given a data table, ReSpark searches for similar-topic reports, parses them into interdependent segments corresponding to analytical objectives, and executes them with new data. It identifies inconsistencies and customizes the objectives, data transformations, and textual descriptions. ReSpark allows users to review real-time outputs, insert new objectives, and modify report content. Its effectiveness was evaluated through comparative and user studies.