Media
The Video Game Industry Is More Successful Than Ever. Why Are Its Workers Treated Like Garbage?
Video game workers--whatever their job, employer, or status--have clearly had enough. This month alone, the labor movement has made some of its biggest advancements ever in organizing the techies, artists, and creatives who keep the largest, most culturally significant sector of the global entertainment industry running and thriving. First, on July 19, came "wall-to-wall" union approval at Fallout-maker Bethesda Game Studios, which meant that everyone from engineers to artists could establish a comprehensive unit with the Communications Workers of America. They quickly earned recognition from parent company Microsoft, marking the first wall-to-wall effort to succeed at any of the Big Tech firm's gaming studios. On July 24, even more company workers got into the game.
Star Wars Outlaws: what to expect from Ubisoft's galactic adventure
About 10 minutes into the latest preview build of Star Wars Outlaws, Ubisoft's forthcoming open-world adventure, lead character Kay Vess enters Mirogana: a densely populated, worn-down city on the desolate moon of Toshara. Around us is a mix of sandstone hovels and metallic sci-fi buildings, crammed with flickering computer panels, neon signs and holographic adverts. Exotic aliens lurk in quiet corners, R2 droids glide past twittering to themselves. Nearby is a cantina, its shady clientele visible through the smoky doorway, and just to the side is a dimly lit gambling parlour. As you explore, robotic voices read out imperial propaganda over public address systems and stormtroopers patrol the streets, checking IDs. At least as far as this lifelong Star Wars fan is concerned, these moments perfectly capture the aesthetics and atmosphere of the original trilogy.
Perplexity will put ads in its AI search engine and share revenue with publishers
When people type a question into Perplexity, the two-year-old search engine scours the internet and uses information from multiple sources, including online publishers, to synthesize an answer using AI. Soon, Perplexity will start sharing revenue with some publishers as part of an advertising platform it plans to launch around the end of September, the company announced on Tuesday. The initiative, known as the Perplexity Publishers' Program, comes less than two months after the San Francisco-based startup backed by investors like Jeff Bezos and NVIDIA, and valued at 3 billion, came under fire from Forbes, Wired, and Condé Nast for allegedly scraping content without permission and ignoring robots.txt, Perplexity's initial partners include TIME, Fortune, The Texas Tribune, Der Spiegel and Automattic, the company behind Wordpress.com. It's not clear exactly how much revenue Perplexity will share with publishers.
Emotion-driven Piano Music Generation via Two-stage Disentanglement and Functional Representation
Huang, Jingyue, Chen, Ke, Yang, Yi-Hsuan
Managing the emotional aspect remains a challenge in automatic music generation. Prior works aim to learn various emotions at once, leading to inadequate modeling. This paper explores the disentanglement of emotions in piano performance generation through a two-stage framework. The first stage focuses on valence modeling of lead sheet, and the second stage addresses arousal modeling by introducing performance-level attributes. To further capture features that shape valence, an aspect less explored by previous approaches, we introduce a novel functional representation of symbolic music. This representation aims to capture the emotional impact of major-minor tonality, as well as the interactions among notes, chords, and key signatures. Objective and subjective experiments validate the effectiveness of our framework in both emotional valence and arousal modeling. We further leverage our framework in a novel application of emotional controls, showing a broad potential in emotion-driven music generation.
Lyrics Transcription for Humans: A Readability-Aware Benchmark
Cífka, Ondřej, Schreiber, Hendrik, Miner, Luke, Stöter, Fabian-Robert
Writing down lyrics for human consumption involves not only accurately capturing word sequences, but also incorporating punctuation and formatting for clarity and to convey contextual information. This includes song structure, emotional emphasis, and contrast between lead and background vocals. While automatic lyrics transcription (ALT) systems have advanced beyond producing unstructured strings of words and are able to draw on wider context, ALT benchmarks have not kept pace and continue to focus exclusively on words. To address this gap, we introduce Jam-ALT, a comprehensive lyrics transcription benchmark. The benchmark features a complete revision of the JamendoLyrics dataset, in adherence to industry standards for lyrics transcription and formatting, along with evaluation metrics designed to capture and assess the lyric-specific nuances, laying the foundation for improving the readability of lyrics. We apply the benchmark to recent transcription systems and present additional error analysis, as well as an experimental comparison with a classical music dataset.
Abstractive summarization from Audio Transcription
Currently, large language models are gaining popularity, their achievements are used in many areas, ranging from text translation to generating answers to queries. However, the main problem with these new machine learning algorithms is that training such models requires large computing resources that only large IT companies have. To avoid this problem, a number of methods (LoRA, quantization) have been proposed so that existing models can be effectively fine-tuned for specific tasks. In this paper, we propose an E2E (end to end) audio summarization model using these techniques. In addition, this paper examines the effectiveness of these approaches to the problem under consideration and draws conclusions about the applicability of these methods.
Enabling Contextual Soft Moderation on Social Media through Contrastive Textual Deviation
Paudel, Pujan, Saeed, Mohammad Hammas, Auger, Rebecca, Wells, Chris, Stringhini, Gianluca
Automated soft moderation systems are unable to ascertain if a post supports or refutes a false claim, resulting in a large number of contextual false positives. This limits their effectiveness, for example undermining trust in health experts by adding warnings to their posts or resorting to vague warnings instead of granular fact-checks, which result in desensitizing users. In this paper, we propose to incorporate stance detection into existing automated soft-moderation pipelines, with the goal of ruling out contextual false positives and providing more precise recommendations for social media content that should receive warnings. We develop a textual deviation task called Contrastive Textual Deviation (CTD) and show that it outperforms existing stance detection approaches when applied to soft moderation.We then integrate CTD into the stateof-the-art system for automated soft moderation Lambretta, showing that our approach can reduce contextual false positives from 20% to 2.1%, providing another important building block towards deploying reliable automated soft moderation tools on social media.
Computational music analysis from first principles
Tymoczko, Dmitri, Newman, Mark
We use coupled hidden Markov models to automatically annotate the 371 Bach chorales in the Riemenschneider edition, a corpus containing approximately 100,000 notes and 20,000 chords. We give three separate analyses that achieve progressively greater accuracy at the cost of making increasingly strong assumptions about musical syntax. Although our method makes almost no use of human input, we are able to identify both chords and keys with an accuracy of 85% or greater when compared to an expert human analysis, resulting in annotations accurate enough to be used for a range of music-theoretical purposes, while also being free of subjective human judgments. Our work bears on longstanding debates about the objective reality of the structures postulated by standard Western harmonic theory, as well as on specific questions about the nature of Western harmonic syntax.
Samsung Galaxy Flip 6 review: A slightly better foldable aimed at everyone
Samsung's Galaxy Z Flip series has always tempted me more than the Z Fold. Maybe it's the flip-phone nostalgia taking hold; maybe it's the fact that I don't want to watch video inside a square; maybe it's simply the Z Flip's more palatable price. The Z Flip series has launched in tandem with the Z Fold for several years, but often with specifications that put it around the bottom of each flagship family, including the traditionally shaped Galaxy S family. As we mentioned in our Z Fold 6 review, there's more foldable competition than ever. In fact, in the face of Motorola's most recent foldables, while Samsung is doing something, is it enough? While Z Flip 6's design has remained largely the same, Samsung made several under-the-hood upgrades this year, with improved battery life and cameras.
ImagiNet: A Multi-Content Dataset for Generalizable Synthetic Image Detection via Contrastive Learning
Boychev, Delyan, Cholakov, Radostin
Generative models, such as diffusion models (DMs), variational autoencoders (VAEs), and generative adversarial networks (GANs), produce images with a level of authenticity that makes them nearly indistinguishable from real photos and artwork. While this capability is beneficial for many industries, the difficulty of identifying synthetic images leaves online media platforms vulnerable to impersonation and misinformation attempts. To support the development of defensive methods, we introduce ImagiNet, a high-resolution and balanced dataset for synthetic image detection, designed to mitigate potential biases in existing resources. It contains 200K examples, spanning four content categories: photos, paintings, faces, and uncategorized. Synthetic images are produced with open-source and proprietary generators, whereas real counterparts of the same content type are collected from public datasets. The structure of ImagiNet allows for a two-track evaluation system: i) classification as real or synthetic and ii) identification of the generative model. To establish a baseline, we train a ResNet-50 model using a self-supervised contrastive objective (SelfCon) for each track. The model demonstrates state-of-the-art performance and high inference speed across established benchmarks, achieving an AUC of up to 0.99 and balanced accuracy ranging from 86% to 95%, even under social network conditions that involve compression and resizing.