Media
Joe Rogan scorches 'liberal robot zombie' phenomenon that brainwashes people: 'Can't think for themselves'
Podcast giant Joe Rogan slammed academia for defending unhealthy lifestyles in a Saturday episode of his "The Joe Rogan Experience" show. Comedians Matt McCusker and Shane Gillis jointed "The Joe Rogan Experience" podcast on Saturday to talk about how media and academia turn people into mindless servants of the establishment. Rogan spoke with the comedians, one of whom had been infamously canceled by Saturday Night Live for past offensive humor. Rogan brought up a viral video showing political performance artist Alex Stein trolling an activist, "This right-wing comedian, he goes to one of these Ukraine protests and he brings a homeless guy, and he says'my wife's boyfriend is homeless, why don't you help him and the homeless people here?'" Rogan followed up by recounting how the activist tried to persuade Stein and the ostensibly homeless man, "The guy, like, tries to engage about the problems with Ukraine," later adding, "this guy is, like, that liberal. That liberal robot zombie repeating sh** that he saw on CNBC, just saying it."
CRYPTEXT: Database and Interactive Toolkit of Human-Written Text Perturbations in the Wild
Le, Thai, Yiran, Ye, Hu, Yifan, Lee, Dongwon
User-generated textual contents on the Internet are often noisy, erroneous, and not in correct forms in grammar. In fact, some online users choose to express their opinions online through carefully perturbed texts, especially in controversial topics (e.g., politics, vaccine mandate) or abusive contexts (e.g., cyberbullying, hate-speech). However, to the best of our knowledge, there is no framework that explores these online ``human-written" perturbations (as opposed to algorithm-generated perturbations). Therefore, we introduce an interactive system called CRYPTEXT. CRYPTEXT is a data-intensive application that provides the users with a database and several tools to extract and interact with human-written perturbations. Specifically, CRYPTEXT helps look up, perturb, and normalize (i.e., de-perturb) texts. CRYPTEXT also provides an interactive interface to monitor and analyze text perturbations online. A short demo video is available at: https://youtu.be/8WT3G8xjIoI
Msanii: High Fidelity Music Synthesis on a Shoestring Budget
Our model combines the expressiveness of mel spectrograms, the generative capabilities of diffusion models, and the vocoding capabilities of neural vocoders. We demonstrate the effectiveness of Msanii by synthesizing tens of seconds (190 seconds) of stereo music at high sample rates (44.1 kHz) without the use of concatenative synthesis, cascading architectures, or compression techniques. To the best of our knowledge, this is the first work to successfully employ a diffusion-based model for synthesizing such long music samples at high sample rates. Our demo can be found here and our code here.
Computational Assessment of Hyperpartisanship in News Titles
Lyu, Hanjia, Pan, Jinsheng, Wang, Zichen, Luo, Jiebo
We first adopt a human-guided machine learning framework to develop a new dataset for hyperpartisan news title detection with 2,200 manually labeled and 1.8 million machine-labeled titles that were posted from 2014 to the present by nine representative media organizations across three media bias groups - Left, Central, and Right in an active learning manner. The fine-tuned transformer-based language model achieves an overall accuracy of 0.84 and an F1 score of 0.78 on an external validation set. Next, we conduct a computational analysis to quantify the extent and dynamics of partisanship in news titles. While some aspects are as expected, our study reveals new or nuanced differences between the three media groups. We find that overall the Right media tends to use proportionally more hyperpartisan titles. Roughly around the 2016 Presidential Election, the proportions of hyperpartisan titles increased in all media bias groups where the relative increase in the proportion of hyperpartisan titles of the Left media was the most. We identify three major topics including foreign issues, political systems, and societal issues that are suggestive of hyperpartisanship in news titles using logistic regression models and the Shapley values. Through an analysis of the topic distribution, we find that societal issues gradually receive more attention from all media groups. We further apply a lexicon-based language analysis tool to the titles of each topic and quantify the linguistic distance between any pairs of the three media groups. Three distinct patterns are discovered. The Left media is linguistically more different from Central and Right in terms of foreign issues. The linguistic distance between the three media groups becomes smaller over recent years. In addition, a seasonal pattern where linguistic difference is associated with elections is observed for societal issues.
Using Kaldi for Automatic Speech Recognition of Conversational Austrian German
Linke, Julian, Wepner, Saskia, Kubin, Gernot, Schuppler, Barbara
As dialogue systems are becoming more and more interactional and social, also the accurate automatic speech recognition (ASR) of conversational speech is of increasing importance. This shifts the focus from short, spontaneous, task-oriented dialogues to the much higher complexity of casual face-to-face conversations. However, the collection and annotation of such conversations is a time-consuming process and data is sparse for this specific speaking style. This paper presents ASR experiments with read and conversational Austrian German as target. In order to deal with having only limited resources available for conversational German and, at the same time, with a large variation among speakers with respect to pronunciation characteristics, we improve a Kaldi-based ASR system by incorporating a (large) knowledge-based pronunciation lexicon, while exploring different data-based methods to restrict the number of pronunciation variants for each lexical entry. We achieve best WER of 0.4% on Austrian German read speech and best average WER of 48.5% on conversational speech. We find that by using our best pronunciation lexicon a similarly high performance can be achieved than by increasing the size of the data used for the language model by approx. 360% to 760%. Our findings indicate that for low-resource scenarios -- despite the general trend in speech technology towards using data-based methods only -- knowledge-based approaches are a successful, efficient method.
Hitting the Books: How to build a music recommendation 'information-space-beast'
As of October, singers, songwriters and music makers are uploading 100,000 new songs every day to streaming services like Spotify. That is too much music. There's no reality, alternate or otherwise, wherein someone could conceivably listen to all that even in a thousand lifetimes. Whether you're into Japanese noise, Russian hardcore, Senegalese afro-house, Swedish doom metal, or Bay Area hip hop, the sheer scale of available listening options is paralyzing. It's a monumental problem that data scientist Glenn McDonald is working to solve.
Bike Frames: Understanding the Implicit Portrayal of Cyclists in the News
Zhao, Xingmeng, Walton, Xavier, Shrestha, Suhana, Rios, Anthony
Increasing the number of cyclists, whether for general transport or recreation, can provide health improvements and reduce the environmental impact of vehicular transportation. However, the public's perception of cycling may be driven by the ideologies and reporting standards of news agencies. For instance, people may identify cyclists on the road as "dangerous" if news agencies overly report cycling accidents, limiting the number of people that cycle for transportation. Moreover, if fewer people cycle, there may be less funding from the government to invest in safe infrastructure. In this paper, we explore the perceived perception of cyclists within news headlines. To accomplish this, we introduce a new dataset, "Bike Frames", that can help provide insight into how headlines portray cyclists and help detect accident-related headlines. Next, we introduce a multi-task (MT) regularization approach that increases the detection accuracy of accident-related posts, demonstrating improvements over traditional MT frameworks. Finally, we compare and contrast the perceptions of cyclists with motorcyclist-related headlines to ground the findings with another related activity for both male- and female-related posts. Our findings show that general news websites are more likely to report accidents about cyclists than other events. Moreover, cyclist-specific websites are more likely to report about accidents than motorcycling-specific websites, even though there is more potential danger for motorcyclists. Finally, we show substantial differences in the reporting about male vs. female-related persons, e.g., more male-related cyclists headlines are related to accidents, but more female-related motorcycling headlines about accidents. WARNING: This paper contains descriptions of accidents and death.