Media
Music Source Separation with Band-split RNN
The performance of music source separation (MSS) models has been greatly improved in recent years thanks to the development of novel neural network architectures and training pipelines. However, recent model designs for MSS were mainly motivated by other audio processing tasks or other research fields, while the intrinsic characteristics and patterns of the music signals were not fully discovered. In this paper, we propose band-split RNN (BSRNN), a frequency-domain model that explictly splits the spectrogram of the mixture into subbands and perform interleaved band-level and sequence-level modeling. The choices of the bandwidths of the subbands can be determined by a priori knowledge or expert knowledge on the characteristics of the target source in order to optimize the performance on a certain type of target musical instrument. To better make use of unlabeled data, we also describe a semi-supervised model finetuning pipeline that can further improve the performance of the model. Experiment results show that BSRNN trained only on MUSDB18-HQ dataset significantly outperforms several top-ranking models in Music Demixing (MDX) Challenge 2021, and the semi-supervised finetuning stage further improves the performance on all four instrument tracks.
Co-Writing Screenplays and Theatre Scripts with Language Models: An Evaluation by Industry Professionals
Mirowski, Piotr, Mathewson, Kory W., Pittman, Jaylen, Evans, Richard
Language models are increasingly attracting interest from writers. However, such models lack long-range semantic coherence, limiting their usefulness for longform creative writing. We address this limitation by applying language models hierarchically, in a system we call Dramatron. By building structural context via prompt chaining, Dramatron can generate coherent scripts and screenplays complete with title, characters, story beats, location descriptions, and dialogue. We illustrate Dramatron's usefulness as an interactive co-creative system with a user study of 15 theatre and film industry professionals. Participants co-wrote theatre scripts and screenplays with Dramatron and engaged in open-ended interviews. We report critical reflections both from our interviewees and from independent reviewers who watched stagings of the works to illustrate how both Dramatron and hierarchical text generation could be useful for human-machine co-creativity. Finally, we discuss the suitability of Dramatron for co-creativity, ethical considerations -- including plagiarism and bias -- and participatory models for the design and deployment of such tools.
Learning by Distilling Context
Snell, Charlie, Klein, Dan, Zhong, Ruiqi
Language models significantly benefit from context tokens, such as prompts or scratchpads. They perform better when prompted with informative instructions, and they acquire new reasoning capabilities by generating a scratch-pad before predicting the final answers. However, they do not \textit{internalize} these performance gains, which disappear when the context tokens are gone. Our work proposes to apply context distillation so that a language model can improve itself by internalizing these gains. Concretely, given a synthetic unlabeled input for the target task, we condition the model on ``[instructions] + [task-input]'' to predict ``[scratch-pad] + [final answer]''; then we fine-tune the same model to predict its own ``[final answer]'' conditioned on the ``[task-input]'', without seeing the ``[instructions]'' or using the ``[scratch-pad]''. We show that context distillation is a general method to train language models, and it can effectively internalize 3 types of training signals. First, it can internalize abstract task instructions and explanations, so we can iteratively update the model parameters with new instructions and overwrite old ones. Second, it can internalize step-by-step reasoning for complex tasks (e.g., 8-digit addition), and such a newly acquired capability proves to be useful for other downstream tasks. Finally, it can internalize concrete training examples, and it outperforms directly learning with gradient descent by 9\% on the SPIDER Text-to-SQL dataset; furthermore, combining context distillation operations can internalize more training examples than the context window size allows.
Graph Attention Network for Camera Relocalization on Dynamic Scenes
Ouali, Mohamed Amine, Bouguessa, Mohamed, Ksantini, Riadh
We devise a graph attention network-based approach for learning a scene triangle mesh representation in order to estimate an image camera position in a dynamic environment. Previous approaches built a scene-dependent model that explicitly or implicitly embeds the structure of the scene. They use convolution neural networks or decision trees to establish 2D/3D-3D correspondences. Such a mapping overfits the target scene and does not generalize well to dynamic changes in the environment. Our work introduces a novel approach to solve the camera relocalization problem by using the available triangle mesh. Our 3D-3D matching framework consists of three blocks: (1) a graph neural network to compute the embedding of mesh vertices, (2) a convolution neural network to compute the embedding of grid cells defined on the RGB-D image, and (3) a neural network model to establish the correspondence between the two embeddings. These three components are trained end-to-end. To predict the final pose, we run the RANSAC algorithm to generate camera pose hypotheses, and we refine the prediction using the point-cloud representation. Our approach significantly improves the camera pose accuracy of the state-of-the-art method from $0.358$ to $0.506$ on the RIO10 benchmark for dynamic indoor camera relocalization.
Dataset Summarization by K Principal Concepts
We propose the new task of K principal concept identification for dataset summarizarion. The objective is to find a set of K concepts that best explain the variation within the dataset. Concepts are high-level human interpretable terms such as "tiger", "kayaking" or "happy". The K concepts are selected from a (potentially long) input list of candidates, which we denote the concept-bank. The concept-bank may be taken from a generic dictionary or constructed by task-specific prior knowledge. An image-language embedding method (e.g. CLIP) is used to map the images and the concept-bank into a shared feature space. To select the K concepts that best explain the data, we formulate our problem as a K-uncapacitated facility location problem. An efficient optimization technique is used to scale the local search algorithm to very large concept-banks. The output of our method is a set of K principal concepts that summarize the dataset. Our approach provides a more explicit summary in comparison to selecting K representative images, which are often ambiguous. As a further application of our method, the K principal concepts can be used to classify the dataset into K groups. Extensive experiments demonstrate the efficacy of our approach.
Physical training is the next hurdle for artificial intelligence, researcher says
Let a million monkeys clack on a million typewriters for a million years and, the adage goes, they'll reproduce the works of Shakespeare. Give infinite monkeys infinite time, and they still will not appreciate the bard's poetic turn-of-phrase, even if they can type out the words. The same holds true for artificial intelligence (AI), according to Michael Woolridge, professor of computer science at the University of Oxford. The issue, he said, is not the processing power, but rather a lack of experience. His perspective was published on July 25 in Intelligent Computing, a Science Partner Journal.
Metatron Inc. Signs Contract to Complete Its First Artificial Intelligence Technology Acquisition
Dover, DE, Sept. 21, 2022 (GLOBE NEWSWIRE) -- Metatron Inc. (OTC Pink: MRNJ), a mobile and web technology pioneer having developed over 2,000 apps on iTunes and Google Play, is pleased to announce that the Company has signed final agreement paperwork with Geek Labs Limited to complete Metatron's first acquisition of artificial intelligence technology. The acquisition was successfully negotiated by the two parties over the course of this past week. The acquisition will be a non-dilutive cash purchase and comes with immediate revenue-generating potential for Metatron. Furthermore, this technology acquisition is a wholly owned asset within the quickly growing $450 billion AI industry and can immediately be added to Metatron's corporate bottom line. CEO Joe Riehl commented: "One week ago we formally announced the formation of Metatron's Artificial Intelligence Division. Today we are announcing that we've inked a deal to complete our first AI tech acquisition. I want to send a crystal-clear message to our valued stakeholders that Metatron is entering a new chapter and taking fast steps to add impressive value for our investors. We already have our eye on a second acquisition and talks with the seller are at the midway point. Furthermore, our new AI division has begun planning projects with talented developers operating within the AI sector to build proprietary Metatron AI technology for use in specialty applications. Our new AI division also has begun performing due diligence on strategic AI patents that our team believes could be used in developing new technology and/or holding big tech companies accountable as more and more AI applications come into use."
Boostr Launches Proposal-IQ, An AI-Based Recommendation Engine for Smarter RFP Responses
Boostr, the only pipeline-to-profits advertising management platform for the media industry announced the launch of Proposal-IQ, a groundbreaking new product designed to help media sales organizations improve the quality of RFP responses while simultaneously building proposals faster. By suggesting the optimal media mix and media plan for each client's objectives and tying into the media seller's real-time inventory, Proposal-IQ drives larger deal sizes, increases in average products sold per order, and higher sell-through--helping media sales organizations grow their revenue. Media companies typically send out dozens of proposals each week and have less than 48 hours to respond to RFPs, creating an enormous burden on media sales teams who need to obtain pricing approvals, conduct inventory checks and complete multiple reviews as part of a highly manual process. Time constraints and repetitiveness often lead salespeople and planners to pitch the same products repeatedly, and their choices are often based more on familiarity than on performance data. Unfortunately, only 37% of proposals contain multiple products yet the top quartile of highest growing publishers are selling multiple products on 48% of proposals.
Why AI Is Eating The Web3 Creator Economy
On August 20, 2011, Marc Andreessen of a16z, published a pivotal story in the Wall Street Journal: Why Software is Eating the World. Today, September 26, 2022, I'm publishing Why AI Is Eating The Web3 Creator Economy. You see, it all stems from the advances in machine and deep learning that has exploded on the scene with DALL-E, MidJourney, Stable Diffusion and now more recently with NVIDIA's announcement about Get3D. NVIDIA's new GET3D AI-powered tool will mess with many recent startups who have developed tools and apps that scan objects to populate metaverse worlds. "Trained using only 2D images, NVIDIA GET3D generates 3D shapes with high-fidelity textures and complex geometric details. These 3D objects are created in the same format used by popular graphics software applications, allowing users to immediately import their shapes into 3D renderers and game engines for further editing."