Goto

Collaborating Authors

 Large Language Model


CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with LLM-Powered Text Description Sourcing and Mining

arXiv.org Artificial Intelligence

--Music similarity retrieval is fundamental for managing and exploring relevant content from large collections in streaming platforms. This paper presents a novel cross-modal contrastive learning framework that leverages the open-ended nature of text descriptions to guide music similarity modeling, addressing the limitations of traditional uni-modal approaches in capturing complex musical relationships. T o overcome the scarcity of high-quality text-music paired data, this paper introduces a dual-source data acquisition approach combining online scraping and LLM-based prompting, where carefully designed prompts leverage LLMs' comprehensive music knowledge to generate contextually rich descriptions. Extensive experiments demonstrate that the proposed framework achieves significant performance improvements over existing benchmarks through objective metrics, subjective evaluations, and real-world A/B testing on the Huawei Music streaming platform. Music similarity retrieval plays an important role in many music information retrieval (MIR) tasks, such as music recommendation [1], personalized playlist generation [2] and background music replacement in video editing [3], [4]. As digital music collections rapidly expand within streaming platforms, accurately identifying similarities between musical pieces has become critical for managing and exploring relevant content from such large collections efficiently.


How Those Studio Ghibli Memes Are a Sign of OpenAI's Trump-Era Shift

TIME - Tech

In one sense, the pivot has been a long time coming. OpenAI began its decade-long life as a research lab that kept its tools under strict lock and key; when it did release early chatbots and image generation models, they had strict content filters that aimed to prevent misuse. But for years it has been widening the accessibility of its tools in an approach it calls "iterative deployment." The release of ChatGPT in November 2022 was the most popular example of this strategy, which the company believes is necessary to help society adapt to the changes AI is bringing. Still, in another sense, the change to OpenAI's model behavior policies has a more recent proximate cause: the 2024 election of President Donald Trump, and the cultural shift that has accompanied the new administration.


Hayao Miyazaki Would Hate You Losers and Your A.I. Slop

Slate

Sign up for the Slatest to get the most insightful analysis, criticism, and advice out there, delivered to your inbox daily. Since OpenAI released an update earlier this week that improved ChatGPT's ability to generate images based on detailed requests, a dark evil has infected the internet, responsible for the shriveling of souls and the wanton destruction of life and nature itself: Studio Ghibli A.I. slop. Social media has been flooded with images of the most random shit imaginable rendered in the signature style of Hayao Miyazaki, the legendary animator and co-founder of the Japanese company Studio Ghibli, renowned for hand-drawn animated films such as Princess Mononoke, Spirited Away, and My Neighbor Totoro. X in particular, Elon Musk's land of the rising bot, is rife with viral posts extolling the virtues of an innovation that steals human-made creations, chews them into paste, and spits out the reassembled remains, stripped of any of the originality, spirit, and labor that makes art art. It's been 24 hours since OpenAI unexpectedly shook the AI image world with 4o image generation.


Anthropic's Claude Is Good at Poetry--and Bullshitting

WIRED

The researchers of Anthropic's interpretability group know that Claude, the company's large language model, is not a human being, or even a conscious piece of software. Still, it's very hard for them to talk about Claude, and advanced LLMs in general, without tumbling down an anthropomorphic sinkhole. Between cautions that a set of digital operations is in no way the same as a cogitating human being, they often talk about what's going on inside Claude's head. It's literally their job to find out. The papers they publish describe behaviors that inevitably court comparisons with real-life organisms.


The Download: peering inside an LLM, and the rise of Signal

MIT Technology Review

April 2024 As the number of satellites in space grows, and as we rely on them for increasing numbers of vital tasks on Earth, the need to better predict stormy space weather is becoming more and more urgent. Scientists have long known that solar activity can change the density of the upper atmosphere. But it's incredibly difficult to precisely predict the sorts of density changes that a given amount of solar activity would produce. Now, experts are working on a model of the upper atmosphere to help scientists to improve their models of how solar activity affects the environment in low Earth orbit. If they succeed, they'll be able to keep satellites safe even amid turbulent space weather, reducing the risk of potentially catastrophic orbital collisions.


If Anthropic Succeeds, a Nation of Benevolent AI Geniuses Could Be Born

WIRED

When Dario Amodei gets excited about AI--which is nearly always--he moves. The cofounder and CEO springs from a seat in a conference room and darts over to a whiteboard. He scrawls charts with swooping hockey-stick curves that show how machine intelligence is bending toward the infinite. His hand rises to his curly mop of hair, as if he's caressing his neurons to forestall a system crash. You can almost feel his bones vibrate as he explains how his company, Anthropic, is unlike other AI model builders.


Copyright questions loom as ChatGPT's Ghibli-style images go viral

The Japan Times

The release of the latest image generator on OpenAI's ChatGPT has triggered a flood of online memes featuring images done in the style of Studio Ghibli, the Japanese studio behind classic animated films like "My Neighbor Totoro" and "Princess Mononoke." Since the release on Wednesday, AI-generated images depicting Studio Ghibli versions of Elon Musk with U.S. President Donald Trump, "The Lord of the Rings," and even a recreation of the Sept. 11 attacks have gone viral across online platforms.


Firm or Fickle? Evaluating Large Language Models Consistency in Sequential Interactions

arXiv.org Artificial Intelligence

Large Language Models (LLMs) have shown remarkable capabilities across various tasks, but their deployment in high-stake domains requires consistent performance across multiple interaction rounds. This paper introduces a comprehensive framework for evaluating and improving LLM response consistency, making three key contributions. First, we propose a novel Position-Weighted Consistency (PWC) score that captures both the importance of early-stage stability and recovery patterns in multi-turn interactions. Second, we present a carefully curated benchmark dataset spanning diverse domains and difficulty levels, specifically designed to evaluate LLM consistency under various challenging follow-up scenarios. Third, we introduce Confidence-Aware Response Generation (CARG), a framework that significantly improves response stability by incorporating model confidence signals into the generation process. Empirical results demonstrate that CARG significantly improves response stability without sacrificing accuracy, underscoring its potential for reliable LLM deployment in critical applications.


Generalization Bias in Large Language Model Summarization of Scientific Research

arXiv.org Artificial Intelligence

Artificial intelligence chatbots driven by large language models (LLMs) have the potential to increase public science literacy and support scientific research, as they can quickly summarize complex scientific information in accessible terms. However, when summarizing scientific texts, LLMs may omit details that limit the scope of research conclusions, leading to generalizations of results broader than warranted by the original study. We tested 10 prominent LLMs, including ChatGPT-4o, ChatGPT-4.5, DeepSeek, LLaMA 3.3 70B, and Claude 3.7 Sonnet, comparing 4900 LLM-generated summaries to their original scientific texts. Even when explicitly prompted for accuracy, most LLMs produced broader generalizations of scientific results than those in the original texts, with DeepSeek, ChatGPT-4o, and LLaMA 3.3 70B overgeneralizing in 26 to 73% of cases. In a direct comparison of LLM-generated and human-authored science summaries, LLM summaries were nearly five times more likely to contain broad generalizations (OR = 4.85, 95% CI [3.06, 7.70]). Notably, newer models tended to perform worse in generalization accuracy than earlier ones. Our results indicate a strong bias in many widely used LLMs towards overgeneralizing scientific conclusions, posing a significant risk of large-scale misinterpretations of research findings. We highlight potential mitigation strategies, including lowering LLM temperature settings and benchmarking LLMs for generalization accuracy.


OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs

arXiv.org Artificial Intelligence

The use of omni-LLMs (large language models that accept any modality as input), particularly for multimodal cognitive state tasks involving speech, is understudied. We present OmniVox, the first systematic evaluation of four omni-LLMs on the zero-shot emotion recognition task. We evaluate on two widely used multimodal emotion benchmarks: IEMOCAP and MELD, and find zero-shot omni-LLMs outperform or are competitive with fine-tuned audio models. Alongside our audio-only evaluation, we also evaluate omni-LLMs on text only and text and audio. We present acoustic prompting, an audio-specific prompting strategy for omni-LLMs which focuses on acoustic feature analysis, conversation context analysis, and step-by-step reasoning. We compare our acoustic prompting to minimal prompting and full chain-of-thought prompting techniques. We perform a context window analysis on IEMOCAP and MELD, and find that using context helps, especially on IEMOCAP. We conclude with an error analysis on the generated acoustic reasoning outputs from the omni-LLMs.