Media
A weapon to surpass Metal Gear - Xe Iaso
Every so often, I like to look at some of the more weird conspiracy theories and then try to debunk them. I consider it a media literacy exercise, but there has been one theory that I've come across that is impressively hard to debunk: the "Dead Internet" theory. I think that the best conspiracy theories are the ones that are hardest to debunk, and this one is increasingly getting more difficult to debunk. The core idea is that the Internet itself is actually dead, no human authorship of any content exists. Any actual human content that is created is isolated into its own little heavenbanned bubble. Mainstream platforms, news outlets, social media sites, Internet forums, chatrooms, everything filled with bot generated content to the point that it's impossible to find another human. To be clear, this theory as literally written is absolute nonsense and probably not worth taking too seriously.
AI in the Artistic Realms. *This is the first part of my…
Keep following to read the upcoming pieces on specific fields like music, poetry, visual arts, and video gaming. Artificial Intelligence (AI) is no longer just a buzzword in the tech industry. Today, AI is making its way into various fields, including art. AI is transforming the way artists create, present, and distribute their work. From music to visual arts, AI is being used in innovative ways to create new experiences for both artists and audiences.
Multimodal Lyrics-Rhythm Matching
Liao, Callie C., Liao, Duoduo, Guessford, Jesse
Despite the recent increase in research on artificial intelligence for music, prominent correlations between key components of lyrics and rhythm such as keywords, stressed syllables, and strong beats are not frequently studied. This is likely due to challenges such as audio misalignment, inaccuracies in syllabic identification, and most importantly, the need for cross-disciplinary knowledge. To address this lack of research, we propose a novel multimodal lyrics-rhythm matching approach in this paper that specifically matches key components of lyrics and music with each other without any language limitations. We use audio instead of sheet music with readily available metadata, which creates more challenges yet increases the application flexibility of our method. Furthermore, our approach creatively generates several patterns involving various multimodalities, including music strong beats, lyrical syllables, auditory changes in a singer's pronunciation, and especially lyrical keywords, which are utilized for matching key lyrical elements with key rhythmic elements. This advantageous approach not only provides a unique way to study auditory lyrics-rhythm correlations including efficient rhythm-based audio alignment algorithms, but also bridges computational linguistics with music as well as music cognition. Our experimental results reveal an 0.81 probability of matching on average, and around 30% of the songs have a probability of 0.9 or higher of keywords landing on strong beats, including 12% of the songs with a perfect landing. Also, the similarity metrics are used to evaluate the correlation between lyrics and rhythm. It shows that nearly 50% of the songs have 0.70 similarity or higher. In conclusion, our approach contributes significantly to the lyrics-rhythm relationship by computationally unveiling insightful correlations.
Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism
Hono, Yukiya, Hashimoto, Kei, Nankaku, Yoshihiko, Tokuda, Keiichi
This paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal modeling is attractive. However, due to the difficulty of the temporal modeling of singing voices, many recent SVS systems with an encoder-decoder-based model still rely on explicitly on duration information generated by additional modules. Although some studies perform simultaneous modeling using seq2seq models with an attention mechanism, they have insufficient robustness against temporal modeling. The proposed attention mechanism is designed to estimate the attention weights by considering the rhythm given by the musical score. Furthermore, several techniques are also introduced to improve the modeling performance of the singing voice. Experimental results indicated that the proposed model is effective in terms of both naturalness and robustness of timing.
FretNet: Continuous-Valued Pitch Contour Streaming for Polyphonic Guitar Tablature Transcription
Cwitkowitz, Frank, Hirvonen, Toni, Klapuri, Anssi
In recent years, the task of Automatic Music Transcription (AMT), whereby various attributes of music notes are estimated from audio, has received increasing attention. At the same time, the related task of Multi-Pitch Estimation (MPE) remains a challenging but necessary component of almost all AMT approaches, even if only implicitly. In the context of AMT, pitch information is typically quantized to the nominal pitches of the Western music scale. Even in more general contexts, MPE systems typically produce pitch predictions with some degree of quantization. In certain applications of AMT, such as Guitar Tablature Transcription (GTT), it is more meaningful to estimate continuous-valued pitch contours. Guitar tablature has the capacity to represent various playing techniques, some of which involve pitch modulation. Contemporary approaches to AMT do not adequately address pitch modulation, and offer only less quantization at the expense of more model complexity. In this paper, we present a GTT formulation that estimates continuous-valued pitch contours, grouping them according to their string and fret of origin. We demonstrate that for this task, the proposed method significantly improves the resolution of MPE and simultaneously yields tablature estimation results competitive with baseline models.
WuYun: Exploring hierarchical skeleton-guided melody generation using knowledge-enhanced deep learning
Zhang, Kejun, Wu, Xinda, Zhang, Tieyao, Huang, Zhijie, Tan, Xu, Liang, Qihao, Wu, Songruoyao, Sun, Lingyun
Although deep learning has revolutionized music generation, existing methods for structured melody generation follow an end-to-end left-to-right note-by-note generative paradigm and treat each note equally. Here, we present WuYun, a knowledge-enhanced deep learning architecture for improving the structure of generated melodies, which first generates the most structurally important notes to construct a melodic skeleton and subsequently infills it with dynamically decorative notes into a full-fledged melody. Specifically, we use music domain knowledge to extract melodic skeletons and employ sequence learning to reconstruct them, which serve as additional knowledge to provide auxiliary guidance for the melody generation process. We demonstrate that WuYun can generate melodies with better long-term structure and musicality and outperforms other state-of-the-art methods by 0.51 on average on all subjective evaluation metrics. Our study provides a multidisciplinary lens to design melodic hierarchical structures and bridge the gap between data-driven and knowledge-based approaches for numerous music generation tasks.
Semantics-enhanced Temporal Graph Networks for Content Popularity Prediction
Zhu, Jianhang, Li, Rongpeng, Chen, Xianfu, Mao, Shiwen, Wu, Jianjun, Zhao, Zhifeng
The surging demand for high-definition video streaming services and large neural network models (e.g., Generative Pre-trained Transformer, GPT) implies a tremendous explosion of Internet traffic. To mitigate the traffic pressure, architectures with in-network storage have been proposed to cache popular contents at devices in closer proximity to users. Correspondingly, in order to maximize caching utilization, it becomes essential to devise an effective popularity prediction method. In that regard, predicting popularity with dynamic graph neural network (DGNN) models achieve remarkable performance. However, DGNN models still suffer from tackling sparse datasets where most users are inactive. Therefore, we propose a reformative temporal graph network, named semantics-enhanced temporal graph network (STGN), which attaches extra semantic information into the user-content bipartite graph and could better leverage implicit relationships behind the superficial topology structure. On top of that, we customize its temporal and structural learning modules to further boost the prediction performance. Specifically, in order to efficiently aggregate the diversified semantics that a content might possess, we design a user-specific attention (UsAttn) mechanism for temporal learning module. Unlike the attention mechanism that only analyzes the influence of genres on content, UsAttn also considers the attraction of semantic information to a specific user. Meanwhile, as for the structural learning, we introduce the concept of positional encoding into our attention-based graph learning and adopt a semantic positional encoding (SPE) function to facilitate the analysis of content-oriented user-association analysis. Finally, extensive simulations verify the superiority of our STGN models and demonstrate the effectiveness in content caching.
Human heuristics for AI-generated language are flawed
Jakesch, Maurice, Hancock, Jeffrey, Naaman, Mor
Human communication is increasingly intermixed with language generated by AI. Across chat, email, and social media, AI systems suggest words, complete sentences, or produce entire conversations. AI-generated language is often not identified as such but presented as language written by humans, raising concerns about novel forms of deception and manipulation. Here, we study how humans discern whether verbal self-presentations, one of the most personal and consequential forms of language, were generated by AI. In six experiments, participants (N = 4,600) were unable to detect self-presentations generated by state-of-the-art AI language models in professional, hospitality, and dating contexts. A computational analysis of language features shows that human judgments of AI-generated language are hindered by intuitive but flawed heuristics such as associating first-person pronouns, use of contractions, or family topics with human-written language. We experimentally demonstrate that these heuristics make human judgment of AI-generated language predictable and manipulable, allowing AI systems to produce text perceived as "more human than human." We discuss solutions, such as AI accents, to reduce the deceptive potential of language generated by AI, limiting the subversion of human intuition.
The AI Apocalypse is Here
Whelp, the apocalypse is upon us. This time the end of the world is brought to you by AI. How else do you explain the unending stream of headlines declaring that AI will eliminate jobs, destroy the education system, and rip the heart and soul out of culture and the arts? What more proof do you need of our imminent demise than that AI is as intelligent as a Wharton MBA? Did you get the panic out of your system? Because AI is also creating incredible opportunities for you, as a leader and innovator, to break through the inertia of the status quo, drive meaningful change, and create enormous value.
Just Nine Out Of 116 AI Professionals In Key Films Are Women, Study Finds - cyberpogo
Report says pattern seen in films such as Ex Machina risks contributing to lack of women in tech. A relentless stream of movies, from Iron Man to Ex Machina, has helped entrench systemic gender inequality in the artificial intelligence industry by portraying AI researchers almost exclusively as men, a study has found. The overwhelming predominance of men as leading AI researchers in movies has shaped public perceptions of the industry, the authors say, and risks contributing to a dramatic lack of women in the tech workforce. Beyond the impact on gender balance, the study raises concerns about the knock-on effects of products that favour male users because they are developed by what the former Microsoft employee Margaret Mitchell called "a sea of dudes". "Given that male engineers have repeatedly been shown to engineer products that are most suitable for and adapted to male users, employing more women is essential for addressing the encoding of bias and pejorative stereotypes into AI technologies," the report's authors write.