Goto

Collaborating Authors

 Media


Hierarchical Generative Modeling of Melodic Vocal Contours in Hindustani Classical Music

arXiv.org Artificial Intelligence

Hindustani music is a performance-driven oral tradition that exhibits the rendition of rich melodic patterns. In this paper, we focus on generative modeling of singers' vocal melodies extracted from audio recordings, as the voice is musically prominent within the tradition. Prior generative work in Hindustani music models melodies as coarse discrete symbols which fails to capture the rich expressive melodic intricacies of singing. Thus, we propose to use a finely quantized pitch contour, as an intermediate representation for hierarchical audio modeling. We propose GaMaDHaNi, a modular two-level hierarchy, consisting of a generative model on pitch contours, and a pitch contour to audio synthesis model. We compare our approach to non-hierarchical audio models and hierarchical models that use a self-supervised intermediate representation, through a listening test and qualitative analysis. We also evaluate audio model's ability to faithfully represent the pitch contour input using Pearson correlation coefficient. By using pitch contours as an intermediate representation, we show that our model may be better equipped to listen and respond to musicians in a human-AI collaborative setting by highlighting two potential interaction use cases (1) primed generation, and (2) coarse pitch conditioning.


Netflix drops a gory new trailer for Terminator Zero, an anime from the studio behind Ghost in the Shell

Engadget

The new Terminator anime heading to Netflix looks absolutely brutal in a trailer that dropped this weekend. Terminator Zero is set in 2022 and 1997 (the year of Judgment Day, as described in Terminator 2) and focuses on new characters: Eiko and the scientist Malcom Lee, who are being hunted by a Terminator. The series was produced by Skydance and Production I.G., the Japanese animation studio behind Ghost in the Shell and Psycho-Pass. Fittingly, it drops on August 29, in a nod to the date of the fictional nuclear annihilation event. You can check out the new trailer below -- but just a heads up for anyone who isn't into anime gore, this clip is packed with it.


RoCP-GNN: Robust Conformal Prediction for Graph Neural Networks in Node-Classification

arXiv.org Machine Learning

Graph Neural Networks (GNNs) have emerged as powerful tools for predicting outcomes in graph-structured data. However, a notable limitation of GNNs is their inability to provide robust uncertainty estimates, which undermines their reliability in contexts where errors are costly. One way to address this issue is by providing prediction sets that contain the true label with a predefined probability margin. Our approach builds upon conformal prediction (CP), a framework that promises to construct statistically robust prediction sets or intervals. There are two primary challenges: first, given dependent data like graphs, it is unclear whether the critical assumption in CP - exchangeability - still holds when applied to node classification. Second, even if the exchangeability assumption is valid for conformalized link prediction, we need to ensure high efficiency, i.e., the resulting prediction set or the interval length is small enough to provide useful information. In this article, we propose a novel approach termed Robust Conformal Prediction for GNNs (RoCP-GNN), which integrates conformal prediction (CP) directly into the GNN training process. This method generates prediction sets, instead of just point predictions, that are valid at a user-defined confidence level, assuming only exchangeability. Our approach robustly predicts outcomes with any predictive GNN model while quantifying the uncertainty in predictions within the realm of graph-based semi-supervised learning (SSL). Experimental results demonstrate that GNN models with size loss provide a statistically significant increase in performance. We validate our approach on standard graph benchmark datasets by coupling it with various state-of-the-art GNNs in node classification. The code will be made available after publication.


How did Donald Trump end up posting Taylor Swift deepfakes?

The Guardian

When Donald Trump shared a slew of AI-generated images this week that falsely depicted Taylor Swift and her fans endorsing his campaign for president, the former US president was amplifying the work of a murky non-profit with aspirations to bankroll rightwing media influencers and a history of spreading misinformation. Several of the images Trump posted on his Truth Social platform, which showed digitally rendered young women in "Swifties for Trump" T-shirts, were the products of the John Milton Freedom Foundation. The group's day-to-day operations appear to revolve around sharing engagement bait on X and seeking millions from donors for a "fellowship program" chaired by a high school sophomore that would award 100,000 to Twitter personalities such as Glenn Greenwald, Andy Ngo and Lara Logan, according to a review of the group's tax records, investor documents and social media output. The John Milton Freedom Foundation did not return request for comment to a set of questions about its operations and fellowship program. After months of retweeting conservative media influencers and echoing Elon Musk's claims that freedom of speech is under attack from leftwing forces, one of the organization's messages found its way to Trump and then his millions of supporters.


Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns

arXiv.org Artificial Intelligence

Any report frames issues to favor a particular interpretation by highlighting or excluding certain aspects of a story. Despite the widespread use of framing in disinformation, framing properties and detection methods remain underexplored outside the English-speaking world. We explore how multilingual framing of the same issue differs systematically. We use eight years of Russia-backed disinformation campaigns, spanning 8k news articles in 4 languages targeting 15 countries. We find that disinformation campaigns consistently and intentionally favor specific framing, depending on the target language of the audience. We further discover how Russian-language articles consistently highlight selected frames depending on the region of the media coverage. We find that the two most prominent models for automatic frame analysis underperform and show high disagreement, highlighting the need for further research.


Uncovering Biases with Reflective Large Language Models

arXiv.org Artificial Intelligence

Biases inherent in human endeavors pose significant challenges for machine learning, particularly in supervised learning that relies on potentially biased "ground truth" data. This reliance, coupled with models' tendency to generalize based on statistical maximal likelihood, can propagate and amplify biases, exacerbating societal issues. To address this, our study proposes a reflective methodology utilizing multiple Large Language Models (LLMs) engaged in a dynamic dialogue to uncover diverse perspectives. By leveraging conditional statistics, information theory, and divergence metrics, this novel approach fosters context-dependent linguistic behaviors, promoting unbiased outputs. Furthermore, it enables measurable progress tracking and explainable remediation actions to address identified biases.


Authorities in northern Iraq report casualties from Turkish drone strike

Al Jazeera

Local authorities and news outlets in northern Iraq's semi-autonomous Kurdish region have said that several people were killed in a Turkish drone strike on Friday, including two journalists. In an initial statement on Friday, the regional authorities said that a car belonging to the Kurdistan Workers' Party (PKK) was struck near the city of Sulaymaniyah, killing a senior PKK official, his guard and his driver. However, a later statement by the Kurdistan regional government's Deputy Prime Minister Qubad Talabani said that the attack targeted a group of journalists, two of whom were killed. "They were two women journalists, not members of an armed force to be a threat to the security and stability of any country or region," Talabani said in a statement. Reporters Without Borders (RSF), a press advocacy organisation, also released a statement denouncing the deaths of the two journalists, identified as 27-year-old Hero Baha'uddin and 40-year-old Golestan Tara from Sterk TV.


This startup wants to be the iTunes of AI content licensing

Engadget

The 28-year-old founders of TollBit, a New York-based startup that is all of six months old, think we're living in the "Napster days" of AI. Just like people of a certain generation downloaded digital music, companies are ripping off vast swaths of the internet without paying the rights holders. They want TollBit to be the iTunes of the AI world. "It's kind of the Wild West right now," Olivia Joslin, the company's co-founder and chief operating officer, told Engadget in an interview. "We want to make it easier for AI companies to pay for the data they need."


Was Linguistic A.I. Created by Accident?

The New Yorker

In the spring of 2017, in a room on the second floor of Google's Building 1965, a college intern named Aidan Gomez stretched out, exhausted. It was three in the morning, and Gomez and Ashish Vaswani, a scientist focussed on natural language processing, were working on their team's contribution to the Neural Information Processing Systems conference, the biggest annual meeting in the field of artificial intelligence. Along with the rest of their eight-person group at Google, they had been pushing flat out for twelve weeks, sometimes sleeping in the office, on couches by a curtain that had a neuron-like pattern. They were nearing the finish line, but Gomez didn't have the energy to go out to a bar and celebrate. He couldn't have even if he'd wanted to: he was only twenty, too young to drink in the United States.


Frequency-aware Feature Fusion for Dense Image Prediction

arXiv.org Artificial Intelligence

Dense image prediction tasks demand features with strong category information and precise spatial boundary details at high resolution. To achieve this, modern hierarchical models often utilize feature fusion, directly adding upsampled coarse features from deep layers and high-resolution features from lower levels. In this paper, we observe rapid variations in fused feature values within objects, resulting in intra-category inconsistency due to disturbed high-frequency features. Additionally, blurred boundaries in fused features lack accurate high frequency, leading to boundary displacement. Building upon these observations, we propose Frequency-Aware Feature Fusion (FreqFusion), integrating an Adaptive Low-Pass Filter (ALPF) generator, an offset generator, and an Adaptive High-Pass Filter (AHPF) generator. The ALPF generator predicts spatially-variant low-pass filters to attenuate high-frequency components within objects, reducing intra-class inconsistency during upsampling. The offset generator refines large inconsistent features and thin boundaries by replacing inconsistent features with more consistent ones through resampling, while the AHPF generator enhances high-frequency detailed boundary information lost during downsampling. Comprehensive visualization and quantitative analysis demonstrate that FreqFusion effectively improves feature consistency and sharpens object boundaries. Extensive experiments across various dense prediction tasks confirm its effectiveness. The code is made publicly available at https://github.com/Linwei-Chen/FreqFusion.