Goto

Collaborating Authors

 Media


Warner settles lawsuit with AI music firm and launches joint venture

BBC News

Warner Music Group (WMG) will begin an artificial intelligence (AI) music venture with technology start-up Suno - a year after it sued the firm in a landmark case. As part of the settlement agreement struck between the two firms, Warner will let users create AI-generated music on Suno using the voices, names and likeness of artists who opt-in to the programme. The record label, which represents artists like Dua Lipa, Coldplay and Ed Sheeran, was among several music giants like Sony Music that sued Suno and a similar platform called Udio. AI-generated content has been controversial, with many artists voicing concerns that it could undermine human songwriters. Starting next year, Suno will roll out new advanced and licensed models to its generative-AI music platform, which allows users to create music based on simple descriptions, said Warner in a statement .


Scientists uncover dark new behavior among bloodthirsty rats that could soon sicken people

Daily Mail - Science & tech

Karoline Leavitt's family member'abruptly arrested' by ICE after living in US for decades Residents in liberal Western US city feel'isolated' as state turns extremely red What HAS happened to Beyoncรฉ? Suddenly desperate, I know what's really going on... and it's ugly: CAROLINE BULLOCK LIZ JONES: Sorry, but it's now time for Kate to stop making excuses'I fell for Joan the moment I saw her': The emotional love letter Sir Richard Branson penned to his'rock' on their anniversary - as he announces her death after 50 years together Ina Garten, 77, vulnerably addresses her decision not to have children: 'I can't imagine my life any other way' Sports broadcaster's wife suffers unimaginable tragedy just before he goes on air New'Hollywood of the South' emerges as booming industry generates $1bn... but long-time residents are furious University of Minnesota program offers guidelines to'reverse the whiteness pandemic' Emmy-winning CBS anchor reveals her devastating health battle: 'I've been silently struggling' Bethany MaGee's family issue heartbreaking statement about her injuries after devout Christian, 26, was set ablaze'by 72-time arrestee' on Chicago train MORE: California squirrels evolving in'shocking' way as scientists investigate key behavioral shift Common rats have learned a shocking and deadly new tactic to kill other animals, which could one day lead to a deadly new pandemic among humans. Scientists witnessed as local brown rats ambushed a colony of bats as they entered two caves in Germany, leaping into the air to catch and kill the nocturnal creatures in droves. Moreover, these rats did this in the middle of the night and without being able to see their surroundings. Researchers from the Leibniz Institute for Evolution and Biodiversity Science said it's the first time common rats have ever been seen in Europe acting with such predatory instincts .


MotionV2V: Editing Motion in a Video

arXiv.org Artificial Intelligence

While generative video models have achieved remarkable fidelity and consistency, applying these capabilities to video editing remains a complex challenge. Recent research has extensively explored motion controllability as a means to enhance text-to-video generation or image animation; however, we identify precise motion control as a promising, yet under-explored, paradigm for editing existing videos. In this work, we propose modifying video motion by directly editing sparse trajectories extracted from the input. W e term the deviation between input and output trajectories a'motion edit' and demonstrate that this representation, when coupled with a generative backbone, enables many powerful video editing capabilities. T o achieve this, we introduce a novel pipeline for generating'motion counterfactuals' -- video pairs that share identical content but distinct motion -- and fine-tune a motion-conditioned video diffusion architecture on this dataset. Our approach allows for edits that start at any timestamp and propagate naturally. In a 4-way head-to-head user study, our model achieves over 65% preference against prior work.


Efficient and Fast Generative-Based Singing Voice Separation using a Latent Diffusion Model

arXiv.org Artificial Intelligence

Extracting individual elements from music mixtures is a valuable tool for music production and practice. While neural networks optimized to mask or transform mixture spectrograms into the individual source(s) have been the leading approach, the source overlap and correlation in music signals poses an inherent challenge. Also, accessing all sources in the mixture is crucial to train these systems, while complicated. Attempts to address these challenges in a generative fashion exist, however, the separation performance and inference efficiency remain limited. In this work, we study the potential of diffusion models to advance toward bridging this gap, focusing on generative singing voice separation relying only on corresponding pairs of isolated vocals and mixtures for training. To align with creative workflows, we leverage latent diffusion: the system generates samples encoded in a compact latent space, and subsequently decodes these into audio. This enables efficient optimization and faster inference. Our system is trained using only open data. We outperform existing generative separation systems, and level the compared non-generative systems on a list of signal quality measures and on interference removal. We provide a noise robustness study on the latent encoder, providing insights on its potential for the task. We release a modular toolkit for further research on the topic.


Generation, Evaluation, and Explanation of Novelists' Styles with Single-Token Prompts

arXiv.org Artificial Intelligence

Abstract--Recent advances in large language models have created new opportunities for stylometry, the study of writing styles and authorship. Two challenges, however, remain central: training generative models when no paired data exist, and evaluating stylistic text without relying only on human judgment. In this work, we present a framework for both generating and evaluating sentences in the style of 19th-century novelists. Large language models are fine-tuned with minimal, single-token prompts to produce text in the voices of authors such as Dickens, Austen, Twain, Alcott, and Melville. T o assess these generative models, we employ a transformer-based detector trained on authentic sentences, using it both as a classifier and as a tool for stylistic explanation. We complement this with syntactic comparisons and explainable AI methods, including attention-based and gradient-based analyses, to identify the linguistic cues that drive stylistic imitation. Our findings show that the generated text reflects the authors' distinctive patterns and that AI-based evaluation offers a reliable alternative to human assessment. All artifacts of this work are published online. The ability to recognize and reproduce an author's writing style has long fascinated both literary scholars and computer scientists. Stylometry, the quantitative study of writing style, rests on the idea that every author leaves behind unconscious patterns in vocabulary, syntax, and rhythm [2, 3]. These patterns have been analyzed for centuries in questions of disputed authorship, the study of literary traditions, and more recently in applications such as security and forensics [4].


From Passive Perception to Active Memory: A Weakly Supervised Image Manipulation Localization Framework Driven by Coarse-Grained Annotations

arXiv.org Artificial Intelligence

Image manipulation localization (IML) faces a fundamental trade-off between minimizing annotation cost and achieving fine-grained localization accuracy. Existing fully-supervised IML methods depend heavily on dense pixel-level mask annotations, which limits scalability to large datasets or real-world deployment. In contrast, the majority of existing weakly-supervised IML approaches are based on image-level labels, which greatly reduce annotation effort but typically lack precise spatial localization. To address this dilemma, we propose BoxPromptIML, a novel weakly-supervised IML framework that effectively balances annotation cost and localization performance. Specifically, we propose a coarse region annotation strategy, which can generate relatively accurate manipulation masks at lower cost. To improve model efficiency and facilitate deployment, we further design an efficient lightweight student model, which learns to perform fine-grained localization through knowledge distillation from a fixed teacher model based on the Segment Anything Model (SAM). Moreover, inspired by the human subconscious memory mechanism, our feature fusion module employs a dual-guidance strategy that actively contextualizes recalled prototypical patterns with real-time observational cues derived from the input. Instead of passive feature extraction, this strategy enables a dynamic process of knowledge recollection, where long-term memory is adapted to the specific context of the current image, significantly enhancing localization accuracy and robustness. Extensive experiments across both in-distribution and out-of-distribution datasets show that Box-PromptIML outperforms or rivals fully-supervised models, while maintaining strong generalization, low annotation cost, and efficient deployment characteristics.


The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models

arXiv.org Artificial Intelligence

Analogical reasoning is at the core of human cognition, serving as an important foundation for a variety of intellectual activities. While prior work has shown that LLMs can represent task patterns and surface-level concepts, it remains unclear whether these models can encode high-level relational concepts and apply them to novel situations through structured comparisons. In this work, we explore this fundamental aspect using proportional and story analogies, and identify three key findings. First, LLMs effectively encode the underlying relationships between analogous entities; both attributive and relational information propagate through mid-upper layers in correct cases, whereas reasoning failures reflect missing relational information within these layers. Second, unlike humans, LLMs often struggle not only when relational information is missing, but also when attempting to apply it to new entities. In such cases, strategically patching hidden representations at critical token positions can facilitate information transfer to a certain extent. Lastly, successful analogical reasoning in LLMs is marked by strong structural alignment between analogous situations, whereas failures often reflect degraded or misplaced alignment. Overall, our findings reveal that LLMs exhibit emerging but limited capabilities in encoding and applying high-level relational concepts, highlighting both parallels and gaps with human cognition.


XiCAD: Camera Activation Detection in the Da Vinci Xi User Interface

arXiv.org Artificial Intelligence

Purpose: Robot-assisted minimally invasive surgery relies on endoscopic video as the sole intraoperative visual feedback. The DaVinci Xi system overlays a graphical user interface (UI) that indicates the state of each robotic arm, including the activation of the endoscope arm. Detecting this activation provides valuable metadata such as camera movement information, which can support downstream surgical data science tasks including tool tracking, skill assessment, or camera control automation. Methods: We developed a lightweight pipeline based on a ResNet18 convolutional neural network to automatically identify the position of the camera tile and its activation state within the DaVinci Xi UI. The model was fine-tuned on manually annotated data from the SurgToolLoc dataset and evaluated across three public datasets comprising over 70,000 frames. Results: The model achieved F1-scores between 0.993 and 1.000 for the binary detection of active cameras and correctly localized the camera tile in all cases without false multiple-camera detections. Conclusion: The proposed pipeline enables reliable, real-time extraction of camera activation metadata from surgical videos, facilitating automated preprocessing and analysis for diverse downstream applications. All code, trained models, and annotations are publicly available.


DUO-TOK: Dual-Track Semantic Music Tokenizer for Vocal-Accompaniment Generation

arXiv.org Artificial Intelligence

Duo-Tok is a source-aware dual-codebook tokenizer for vocal-accompaniment music that targets the growing tension between reconstruction quality and language-model (LM) learnability in modern lyrics-to-song systems. Existing codecs either prioritize high-fidelity reconstruction with difficult-to-model acoustic tokens or compress aggressively into semantic tokens that are LM-friendly but lossy, and they rarely make the tokenizer itself aware of dual-track structure. Duo-Tok follows a four-stage, SSL-centered pipeline: we first pretrain a BEST-RQ-style encoder on large-scale audio, then stabilize and factorize the representation with Gaussian replacement noise and multi-task supervision, before freezing the encoder to learn SimVQ-based dual codebooks with hard routing for vocals and accompaniment, and finally training latent diffusion decoders on top of the discrete tokens. Duo-Tok at 0.75 kbps shifts the empirical reconstruction-generation Pareto frontier, achieving the best music-tagging AP and the lowest vocabulary-normalized LM perplexity among compared codecs while maintaining reconstruction quality comparable to state-of-the-art music tokenizers.


Energy Costs and Neural Complexity Evolution in Changing Environments

arXiv.org Artificial Intelligence

The Cognitive Buffer Hypothesis (CBH) posits that larger brains evolved to enhance survival in changing conditions. However, larger brains also carry higher energy demands, imposing additional metabolic burdens. Alongside brain size, brain organization plays a key role in cognitive ability and, with suitable architectures, may help mitigate energy challenges. This study evolves Artificial Neural Networks (ANNs) used by Reinforcement Learning (RL) agents to investigate how environmental variability and energy costs influence the evolution of neural complexity, defined in terms of ANN size and structure. Results indicate that under energy constraints, increasing seasonality led to smaller ANNs. This challenges CBH and supports the Expensive Brain Hypothesis (EBH), as highly seasonal environments reduced net energy intake and thereby constrained brain size. ANN structural complexity primarily emerged as a byproduct of size, where energy costs promoted the evolution of more efficient networks.