Media
New Mexico Is a Great Place for Sci-Fi
Melinda Snodgrass is the novelist and screenwriter best known for her classic Star Trek: The Next Generation script "The Measure of a Man." Her latest novel, Lucifer's War, pits an unlikely band of heroes against a horde of Lovecraftian monsters that have been spreading fear and ignorance throughout human history. "It's unbelievable now, the kind of nonsense people are accepting, that's being pushed on them by social media," Snodgrass says in Episode 529 of the Geek's Guide to the Galaxy podcast. "I really wanted to make a stand for science and rationality, as opposed to magic and superstition." The book is set in Snodgrass' home state of New Mexico, a place where science and superstition clash in a particularly striking way. "It's a very weird place, where you have Los Alamos laboratory, Sandia laboratories, high-tech, high-energy centers," Snodgrass says, "Some of the finest scientific minds in the world come here to lecture and study and commune with each other, and then on the other side you have people who will balance your aura and sell you a crystal to deal with your cancer."
The Future of Podcasting is AI
Roughly speaking, about 22,000 new podcasts are launched in a month. There are close to 2.5 million (more than 71 million episodes) in the Apple Podcasts directory right now, according to Podcast Industry Insights. And those are just the ones we know about. They're going direct to their listeners, selling premium content and having big success," says Andy Taylor, formerly of BBC Radio and founder of Cardiff-based R&D consultancy Bwlb. And that's to say nothing of the growing volume of podcast-like content, whether created by brands for promotion or event producers that want, for example, to make talks available on-demand. Every piece of content needs to be produced and distributed, whether by audio professionals or folks learning the craft. Therefore, the more they can automate large swaths of production, the more they can focus on the content. "The different places audio is being published have just exploded," explains Jonathan Wyner chief engineer at M Works Mastering and a professor at Berklee College of Music in Boston. "With all those contexts, there is a real motivation and imperative for creators to be more versatile." Not to mention, more productive and efficient. Artificial intelligence (AI) -- software that can automate tasks previously done by humans -- holds the key to handling the tsunami of podcast content. Not only can AI speed up production, it can make podcasts sound better and set the stage for the audio experiences of tomorrow. "AI basically helps take care of repetitive tasks to quicken the workflow of the podcaster," explains Manos Chourdakis, research engineer at Nomono, which develops AI-based podcasting tools. "For example, with AI, you don't have to listen to a whole podcast to find where someone said something wrong, then replace or remove it.
Google to Roll Out App for AI-Generated Artwork, Complicating Copyright Worries
A new Google feature will let consumers use artificial intelligence to bring their fantastical creations to (digital) life by just typing a few words. The app, which Bloomberg reported Thursday is currently under development, will have two functions: users can construct cities with its "City Dreamer" function, or customize a family-friendly cartoon monster with its "Wobble" feature. The tools will be available through Google's AI Test Kitchen app, Douglas Eck, a lead scientist at Google, said at the company's AI@ event in New York on Wednesday. The release date for the new app has not yet been announced. The features will use AI imaging technologies to generate hyper-specific images from even short text descriptions.
Artificial intelligence makes enzyme engineering easy
We revealed that the redox cofactor preference of malic enzymes can be strikingly converted by applying phylogenetic analysis to machine learning, without experimental screening. This method can predict mutation positions and candidate amino acids that affect substrate specificity, which is challenging to infer solely from the crystal structures. Machine learning uses the structurally homologous but functionally distinct enzymes' amino acid sequences as input datasets to efficiently navigate toward the target function, and potentially provide new fundamental insights into enzymeโsubstrate specificity. Osaka, Japan โ You can't expect a pharmaceutical scientist to switch labs to the facilities available in a television studio and expect the same research output. Enzymes behave exactly the same.
The Future of Robotics in the Smart City: A Preview for 2023
TORONTO, Nov. 3, 2022 /CNW/ - Robots, the word derives from a 1920's play but represents a variety of automated forms today. Approximately 2.25 million robotic types worldwide, and over 200 million are predicted by 2030. The future of robotics is assured as we become more familiar with and dependent on them. Robots were first created in Japan, bringing what was science fiction to reality. Singapore, South Korea, China, and Germany are the highest producers of robotics.
Large Scale Radio Frequency Wideband Signal Detection & Recognition
Boegner, Luke, Vanhoy, Garrett, Vallance, Phillip, Gulati, Manbir, Feitzinger, Dresden, Comar, Bradley, Miller, Robert D.
Applications of deep learning to the radio frequency (RF) domain have largely concentrated on the task of narrowband signal classification after the signals of interest have already been detected and extracted from a wideband capture. To encourage broader research with wideband operations, we introduce the WidebandSig53 (WBSig53) dataset which consists of 550 thousand synthetically-generated samples from 53 different signal classes containing approximately 2 million unique signals. We extend the TorchSig signal processing machine learning toolkit for open-source and customizable generation, augmentation, and processing of the WBSig53 dataset. We conduct experiments using state of the art (SoTA) convolutional neural networks and transformers with the WBSig53 dataset. We investigate the performance of signal detection tasks, i.e. detect the presence, time, and frequency of all signals present in the input data, as well as the performance of signal recognition tasks, where networks detect the presence, time, frequency, and modulation family of all signals present in the input data. Two main approaches to these tasks are evaluated with segmentation networks and object detection networks operating on complex input spectrograms. Finally, we conduct comparative analysis of the various approaches in terms of the networks' mean average precision, mean average recall, and the speed of inference.
Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis
Qiu, Changyuan, Wu, Winston, Zhang, Xinliang Frederick, Wang, Lu
Prior work on ideology prediction has largely focused on single modalities, i.e., text or images. In this work, we introduce the task of multimodal ideology prediction, where a model predicts binary or five-point scale ideological leanings, given a text-image pair with political content. We first collect five new large-scale datasets with English documents and images along with their ideological leanings, covering news articles from a wide range of US mainstream media and social media posts from Reddit and Twitter. We conduct in-depth analyses of news articles and reveal differences in image content and usage across the political spectrum. Furthermore, we perform extensive experiments and ablation studies, demonstrating the effectiveness of targeted pretraining objectives on different model components. Our best-performing model, a late-fusion architecture pretrained with a triplet objective over multimodal content, outperforms the state-of-the-art text-only model by almost 4% and a strong multimodal baseline with no pretraining by over 3%.
GlobalFlowNet: Video Stabilization using Deep Distilled Global Motion Estimates
James, Jerin Geo, Jain, Devansh, Rajwade, Ajit
Videos shot by laymen using hand-held cameras contain undesirable shaky motion. Estimating the global motion between successive frames, in a manner not influenced by moving objects, is central to many video stabilization techniques, but poses significant challenges. A large body of work uses 2D affine transformations or homography for the global motion. However, in this work, we introduce a more general representation scheme, which adapts any existing optical flow network to ignore the moving objects and obtain a spatially smooth approximation of the global motion between video frames. We achieve this by a knowledge distillation approach, where we first introduce a low pass filter module into the optical flow network to constrain the predicted optical flow to be spatially smooth. This becomes our student network, named as \textsc{GlobalFlowNet}. Then, using the original optical flow network as the teacher network, we train the student network using a robust loss function. Given a trained \textsc{GlobalFlowNet}, we stabilize videos using a two stage process. In the first stage, we correct the instability in affine parameters using a quadratic programming approach constrained by a user-specified cropping limit to control loss of field of view. In the second stage, we stabilize the video further by smoothing global motion parameters, expressed using a small number of discrete cosine transform coefficients. In extensive experiments on a variety of different videos, our technique outperforms state of the art techniques in terms of subjective quality and different quantitative measures of video stability. The source code is publicly available at \href{https://github.com/GlobalFlowNet/GlobalFlowNet}{https://github.com/GlobalFlowNet/GlobalFlowNet}
The (In)Effectiveness of Intermediate Task Training For Domain Adaptation and Cross-Lingual Transfer Learning
Mohapatra, Sovesh, Mohapatra, Somesh
Transfer learning from large language models (LLMs) has emerged as a powerful technique to enable knowledge-based fine-tuning for a number of tasks, adaptation of models for different domains and even languages. However, it remains an open question, if and when transfer learning will work, i.e. leading to positive or negative transfer. In this paper, we analyze the knowledge transfer across three natural language processing (NLP) tasks - text classification, sentimental analysis, and sentence similarity, using three LLMs - BERT, RoBERTa, and XLNet - and analyzing their performance, by fine-tuning on target datasets for domain and cross-lingual adaptation tasks, with and without an intermediate task training on a larger dataset. Our experiments showed that fine-tuning without an intermediate task training can lead to a better performance for most tasks, while more generalized tasks might necessitate a preceding intermediate task training step. We hope that this work will act as a guide on transfer learning to NLP practitioners.