Media
Re-creation of Creations: A New Paradigm for Lyric-to-Melody Generation
Lv, Ang, Tan, Xu, Qin, Tao, Liu, Tie-Yan, Yan, Rui
Lyric-to-melody generation is an important task in songwriting, and is also quite challenging due to its unique characteristics: the generated melodies should not only follow good musical patterns, but also align with features in lyrics such as rhythms and structures. These characteristics cannot be well handled by neural generation models that learn lyric-to-melody mapping in an end-to-end way, due to several issues: (1) lack of aligned lyric-melody training data to sufficiently learn lyric-melody feature alignment; (2) lack of controllability in generation to better and explicitly align the lyric-melody features. In this paper, we propose Re-creation of Creations (ROC), a new paradigm for lyric-to-melody generation. ROC generates melodies according to given lyrics and also conditions on user-designated chord progression. It addresses the above issues through a generation-retrieval pipeline. Specifically, our paradigm has two stages: (1) creation stage, where a huge amount of music fragments generated by a neural melody language model are indexed in a database through several key features (e.g., chords, tonality, rhythm, and structural information); (2) re-creation stage, where melodies are re-created by retrieving music fragments from the database according to the key features from lyrics and concatenating best music fragments based on composition guidelines and melody language model scores. ROC has several advantages: (1) It only needs unpaired melody data to train melody language model, instead of paired lyric-melody data in previous models. (2) It achieves good lyric-melody feature alignment in lyric-to-melody generation. Tested by English and Chinese lyrics, ROC outperforms previous neural based lyric-to-melody generation models on both objective and subjective metrics.
Amazon is reportedly making a Tomb Raider TV series
Hollywood may be taking another stab at a Tomb Raider production, but this time for the small screen. The Hollywood Reporter sources say Amazon is creating a Tomb Raider TV series for Prime Video, with Phoebe Waller-Bridge (of Fleabag fame) set to be an executive producer and write the script. It's not certain who would star, but we wouldn't count on movie stars Angelina Jolie or Alicia Vikander reprising the role of Lara Croft. The show is reportedly still in the development stage. We've asked Amazon for comment.
Phoebe Waller-Bridge reportedly writing Tomb Raider TV series
Phoebe Waller-Bridge is reportedly set to write a new take on Tomb Raider for Amazon. According to the Hollywood Reporter, sources claim the Emmy-winning star and creator of Fleabag is developing a new TV series based on the popular game, writing scripts and executive producing. Amazon is yet to confirm the news. The character of Lara Croft, an archaeological adventurer, has previously been brought to the big screen by Angelina Jolie in two films and more recently, Alicia Vikander in 2018. After that film underperformed, a planned sequel was cancelled and in 2022, MGM lost their rights to the franchise.
Stunning drone footage captures a huge pod of dolphins off the coast of Florida
An armature drone photographer captured stunning footage of a dolphin pod swimming through the crystal-blue waters off the coast of Florida. Local restaurant owner Paul Dabill, 48, filmed approximately 50 dolphins while'looking for life to film' around Jupiter last week. The mesmerizing video shows the marine animals diving in and out of the sea and playing keep-away with a strand of sargassum seaweed. Dabill said he spent 30 minutes filing the pod, one of the largest he had seen. The clip was captured on January 18, when the skies were clear and the ocean was blue.
World champion 'speedcuber' claims the violin has aided in his success with Rubik's Cubes
Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. A University of Michigan student is one of the world's foremost "speedcubers," a person capable of quickly solving a Rubik's Cube. He also is an accomplished violinist. Stanley Chapel says the two fields go hand in hand.
Down the Rabbit Hole: Detecting Online Extremism, Radicalisation, and Politicised Hate Speech
Govers, Jarod, Feldman, Philip, Dant, Aaron, Patros, Panos
Social media is a modern person's digital voice to project and engage with new ideas and mobilise communities $\unicode{x2013}$ a power shared with extremists. Given the societal risks of unvetted content-moderating algorithms for Extremism, Radicalisation, and Hate speech (ERH) detection, responsible software engineering must understand the who, what, when, where, and why such models are necessary to protect user safety and free expression. Hence, we propose and examine the unique research field of ERH context mining to unify disjoint studies. Specifically, we evaluate the start-to-finish design process from socio-technical definition-building and dataset collection strategies to technical algorithm design and performance. Our 2015-2021 51-study Systematic Literature Review (SLR) provides the first cross-examination of textual, network, and visual approaches to detecting extremist affiliation, hateful content, and radicalisation towards groups and movements. We identify consensus-driven ERH definitions and propose solutions to existing ideological and geographic biases, particularly due to the lack of research in Oceania/Australasia. Our hybridised investigation on Natural Language Processing, Community Detection, and visual-text models demonstrates the dominating performance of textual transformer-based algorithms. We conclude with vital recommendations for ERH context mining researchers and propose an uptake roadmap with guidelines for researchers, industries, and governments to enable a safer cyberspace.
Leveraging the Third Dimension in Contrastive Learning
Aithal, Sumukh, Goyal, Anirudh, Lamb, Alex, Bengio, Yoshua, Mozer, Michael
Self-Supervised Learning (SSL) methods operate on unlabeled data to learn robust representations useful for downstream tasks. Most SSL methods rely on augmentations obtained by transforming the 2D image pixel map. These augmentations ignore the fact that biological vision takes place in an immersive three-dimensional, temporally contiguous environment, and that low-level biological vision relies heavily on depth cues. Using a signal provided by a pretrained state-of-the-art monocular RGB-to-depth model (the Depth Prediction Transformer, Ranftl et al., 2021), we explore two distinct approaches to incorporating depth signals into the SSL framework. First, we evaluate contrastive learning using an RGB+depth input representation. Second, we use the depth signal to generate novel views from slightly different camera positions, thereby producing a 3D augmentation for contrastive learning. We evaluate these two approaches on three different SSL methods--BYOL, SimSiam, and SwAV--using ImageNette (10 class subset of ImageNet), ImageNet-100 and ImageNet-1k datasets. We find that both approaches to incorporating depth signals improve the robustness and generalization of the baseline SSL methods, though the first approach (with depth-channel concatenation) is superior. For instance, BYOL with the additional depth channel leads to an increase in downstream classification accuracy from 85.3% to 88.0% on ImageNette and 84.1% to 87.0% on ImageNet-C. Biological vision systems evolved in and interact with a three-dimensional world. As an individual moves through the environment, the relative distance of objects is indicated by rich signals extracted by the visual system, from motion parallax to binocular disparity to occlusion cues. These signals play a role in early development to bootstrap an infant's ability to perceive objects in visual scenes (Spelke, 1990; Spelke & Kinzler, 2007) and to reason about physical interactions between objects (Baillargeon, 2004). In the mature visual system, features predictive of occlusion and three-dimensional structure are extracted early and in parallel in the visual processing stream (Enns & Rensink, 1990; 1991), and early vision uses monocular cues to rapidly complete partially-occluded objects (Rensink & Enns, 1998) and binocular cues to guide attention (Nakayama & Silverman, 1986). In short, biological vision systems are designed to leverage the three-dimensional structure of the environment. In contrast, machine vision systems typically consider a 2D RGB image or a sequence of 2D RGB frames to be the relevant signal.
Truth Machines: Synthesizing Veracity in AI Language Models
Munn, Luke, Magee, Liam, Arora, Vanicka
University of Stirling, United Kingdom vanicka.arora@stir.ac.uk Abstract As AI technologies are rolled out into healthcare, academia, human resources, law, and a multitude of other domains, they become de-facto arbiters of truth. But truth is highly contested, with many different definitions and approaches. It then investigates the production of truth in InstructGPT, a large language model, highlighting how data harvesting, model architectures, and social feedback mechanisms weave together disparate understandings of veracity. It conceptualizes this performance as an operationalization of truth, where distinct, often conflicting claims are smoothly synthesized and confidently presented into truth-statements. We argue that these same logics and inconsistencies play out in Instruct's successor, ChatGPT, reiterating truth as a non-trivial problem. We suggest that enriching sociality and thickening "reality" are two promising vectors for enhancing the truth-evaluating capacities of future language models. We conclude, however, by stepping back to consider AI truth-telling as a social practice: what kind of "truth" do we as listeners desire? OpenAI's latest language model appeared to We stress then that truth in AI is not just technical but be powerful and almost magical, generating news articles, also social, cultural, and political, drawing on particular writing poetry, and explaining arcane concepts norms and values. But a week later, the coding the technical matters: translating truth theories into site StackOverflow banned all answers produced actionable architectures and processes updates them by the model. "The primary problem," explained in significant ways. These disparate sociotechnical the staff, "is that while the answers which ChatGPT forces coalesce into a final AI model which purports produces have a high rate of being incorrect, they typically to tell the truth--and in doing so, our understanding look like they might be good and the answers of "truth" is remade. "The ideal of truth is a fallacy are very easy to produce" (Vincent 2022). For a site for semantic interpretation and needs to be changed," aiming to provide correct answers to coding problems, suggested two AI researchers (Welty and Aroyo 2015).
Madison Square Garden CEO James Dolan threatens to stop alcohol sales at Rangers game
Kelly Conlon, who was kept from seeing the Rockettes, and Sam Davis, who was barred from attending a Rangers game, speak out against MSG Entertainment and James Dolan for their use of facial recognition on'America's Newsroom.' The latest development to come from Madison Square Garden and CEO James Dolan is one that will likely leave fans very unhappy. Dolan threatened to cancel all alcohol sales at The Garden – he mentioned a Rangers game – as a response to the New York State Liquor Authority, which is currently investigating Dolan regarding his facial recognition technology that has resulted in several bans against lawyers who are suing him. Dolan said it all on Fox 5's "Good Day New York." with Rosanna Scotto. James Dolan, left, and head coach Tom Thibodeau of the New York Knicks attend the NBA Summer League at the Thomas and Mack Center on July 8, 2022, in Las Vegas.
Then call them 'robots' • TechCrunch
Before they were robots, they were "androids" or "automatons." The word "robot" is commonly accepted as having arrived in English through -- of all places -- a Czech play. "R.U.R." made its public debut in Prague 102 years ago, yesterday. It would arrive in the States a year and a half later, with Spencer Tracy making his nonspeaking Broadway debut as one of Rossum's titular Universal Robots. The playwright Karel Čapek humbly noted the following decade that he couldn't take full credit for the word's origin.