Goto

Collaborating Authors

 Media


Mesostructures: Beyond Spectrogram Loss in Differentiable Time-Frequency Analysis

arXiv.org Artificial Intelligence

Computer musicians refer to mesostructures as the intermediate levels of articulation between the microstructure of waveshapes and the macrostructure of musical forms. Examples of mesostructures include melody, arpeggios, syncopation, polyphonic grouping, and textural contrast. Despite their central role in musical expression, they have received limited attention in deep learning. Currently, autoencoders and neural audio synthesizers are only trained and evaluated at the scale of microstructure: i.e., local amplitude variations up to 100 milliseconds or so. In this paper, we formulate and address the problem of mesostructural audio modeling via a composition of a differentiable arpeggiator and time-frequency scattering. We empirically demonstrate that time--frequency scattering serves as a differentiable model of similarity between synthesis parameters that govern mesostructure. By exposing the sensitivity of short-time spectral distances to time alignment, we motivate the need for a time-invariant and multiscale differentiable time--frequency model of similarity at the level of both local spectra and spectrotemporal modulations.


XNLI: Explaining and Diagnosing NLI-based Visual Data Analysis

arXiv.org Artificial Intelligence

Natural language interfaces (NLIs) enable users to flexibly specify analytical intentions in data visualization. However, diagnosing the visualization results without understanding the underlying generation process is challenging. Our research explores how to provide explanations for NLIs to help users locate the problems and further revise the queries. We present XNLI, an explainable NLI system for visual data analysis. The system introduces a Provenance Generator to reveal the detailed process of visual transformations, a suite of interactive widgets to support error adjustments, and a Hint Generator to provide query revision hints based on the analysis of user queries and interactions. Two usage scenarios of XNLI and a user study verify the effectiveness and usability of the system. Results suggest that XNLI can significantly enhance task accuracy without interrupting the NLI-based analysis process.


Automated Identification of Disaster News For Crisis Management Using Machine Learning

arXiv.org Artificial Intelligence

A lot of news sources picked up on Typhoon Rai (also known locally as Typhoon Odette), along with fake news outlets. The study honed in on the issue, to create a model that can identify between legitimate and illegitimate news articles. With this in mind, we chose the following machine learning algorithms in our development: Logistic Regression, Random Forest and Multinomial Naive Bayes. Bag of Words, TF-IDF and Lemmatization were implemented in the Model. Gathering 160 datasets from legitimate and illegitimate sources, the machine learning was trained and tested. By combining all the machine learning techniques, the Combined BOW model was able to reach an accuracy of 91.07%, precision of 88.33%, recall of 94.64%, and F1 score of 91.38% and Combined TF-IDF model was able to reach an accuracy of 91.18%, precision of 86.89%, recall of 94.64%, and F1 score of 90.60%.


E-NeRF: Neural Radiance Fields from a Moving Event Camera

arXiv.org Artificial Intelligence

Estimating neural radiance fields (NeRFs) from "ideal" images has been extensively studied in the computer vision community. Most approaches assume optimal illumination and slow camera motion. These assumptions are often violated in robotic applications, where images may contain motion blur, and the scene may not have suitable illumination. This can cause significant problems for downstream tasks such as navigation, inspection, or visualization of the scene. To alleviate these problems, we present E-NeRF, the first method which estimates a volumetric scene representation in the form of a NeRF from a fast-moving event camera. Our method can recover NeRFs during very fast motion and in high-dynamic-range conditions where frame-based approaches fail. We show that rendering high-quality frames is possible by only providing an event stream as input. Furthermore, by combining events and frames, we can estimate NeRFs of higher quality than state-of-the-art approaches under severe motion blur. We also show that combining events and frames can overcome failure cases of NeRF estimation in scenarios where only a few input views are available without requiring additional regularization.


Same Words, Different Meanings: Semantic Polarization in Broadcast Media Language Forecasts Polarization on Social Media Discourse

arXiv.org Artificial Intelligence

With the growth of online news over the past decade, empirical studies on political discourse and news consumption have focused on the phenomenon of filter bubbles and echo chambers. Yet recently, scholars have revealed limited evidence around the impact of such phenomenon, leading some to argue that partisan segregation across news audiences cannot be fully explained by online news consumption alone and that the role of traditional legacy media may be as salient in polarizing public discourse around current events. In this work, we expand the scope of analysis to include both online and more traditional media by investigating the relationship between broadcast news media language and social media discourse. By analyzing a decade's worth of closed captions (2 million speaker turns) from CNN and Fox News along with topically corresponding discourse from Twitter, we provide a novel framework for measuring semantic polarization between America's two major broadcast networks to demonstrate how semantic polarization between these outlets has evolved (Study 1), peaked (Study 2) and influenced partisan discussions on Twitter (Study 3) across the last decade. Our results demonstrate a sharp increase in polarization in how topically important keywords are discussed between the two channels, especially after 2016, with overall highest peaks occurring in 2020. The two stations discuss identical topics in drastically distinct contexts in 2020, to the extent that there is barely any linguistic overlap in how identical keywords are contextually discussed. Further, we demonstrate at scale, how such partisan division in broadcast media language significantly shapes semantic polarity trends on Twitter (and vice-versa), empirically linking for the first time, how online discussions are influenced by televised media.


Innovation and the Pandemic Propelled Performance G.R. Jenkin & Associ

#artificialintelligence

Innovation and the Pandemic Propelled Performance The 2022 TMT Value Creators Report February 28, 2022 By Simon Bamberger, Hady Farag, Derek Kennedy, Franck Luisada, Michaela Novakov, Vaishali Rastogi, and Neal Zuckerman The outsized role that technology, media, and telecommunications (TMT) companies play in modern life has made the sector a leader in creating shareholder value. From 2016 to 2021, TMT companies collectively outperformed those in many other industries in total shareholder return (TSR), according to BCG's 2022 Value Creators Report. Among our findings: The lion's share of TMT value creation came from tech players, which from 2016 to 2021 had a median annual TSR performance of 30%, more than double the median overall return of 13% for the 33 industries we studied. Of the 232 TMT companies in our sample, 70% posted a higher TSR during the period that included the peak of the pandemic, the 21 months from March 2020 to November 2021, than during the prior 21 months. Continuing on the same growth trajectory may be a challenge given recent investor anxieties about inflation, monetary policy, and moderating earnings growth.


How to stop facial recognition cameras from monitoring your every move

FOX News

Apple's got a new helpful feature called "Safety Check" that'll guide you through what you've shared, with whom and how to revoke access. If you ever felt like someone was tracking you, be sure to review these settings. Are you concerned about facial recognition cameras monitoring your every move? Some large venues and arenas are using it as a security measure, claiming it ensures safety for guests and employees. However, the technology is also being used for surveillance and to block people from entering businesses.


UG2+ Challenge

#artificialintelligence

The rapid development of computer vision algorithms increasingly allows automatic visual recognition to be incorporated into a suite of emerging applications. Some of these applications have less-than-ideal circumstances such as low-visibility environments, causing image captures to have degradations. In other more extreme applications, such as imagers for flexible wearables, smart clothing sensors, ultra-thin headset cameras, implantable in vivo imaging, and others, standard camera systems cannot even be deployed, requiring new types of imaging devices. Computational photography addresses the concerns above by designing new computational techniques and incorporating them into the image capture and formation pipeline. This raises a set of new questions.


Real-life insurance lawyer assesses damage in superhero movies and TV shows

#artificialintelligence

Florida-based Stacey Giulianti is both an insurance lawyer and a comic book fan. Who better to review and assess the damage in superhero movies and television shows, such as Man of Steel, The Batman, Spider-Man: Homecoming, 'Venom, The Boys and Avengers: Infinity War, than a 30-year veteran in the field? https://youtu.be/QN7rlarPF6I


DALL·E 2, Explained: The Promise And Limitations Of A Revolutionary AI

#artificialintelligence

DALL·E 2 is the newest AI model by OpenAI. If you've seen some of its creations and think they're amazing, keep reading to understand why you're totally right -- but also wrong. OpenAI published a blog post and a paper entitled "Hierarchical Text-Conditional Image Generation with CLIP Latents" on DALL·E 2. The post is fine if you want to get a glimpse at the results and the paper is great for understanding the technical details, but neither explains DALL·E 2's amazingness -- and the not-so-amazing -- in depth. That's what this article is for. If this in-depth educational content is useful for you, subscribe to our AI mailing list to be alerted when we release new material. DALL·E 2 is the new version of DALL·E, a generative language model that takes sentences and creates corresponding original images. At 3.5B parameters, DALL·E 2 is a large model but not nearly as large as GPT-3 and, interestingly, smaller than its predecessor (12B). Despite its size, DALL·E 2 generates 4x better resolution images than DALL·E and it's preferred by human judges 70% of the time both in caption matching and photorealism. As they did with DALL·E, OpenAI didn't release DALL·E 2 (you can always join the never-ending waitlist). However, they open-sourced CLIP which, although only indirectly related to DALL·E, forms the basis of DALL·E 2. (CLIP is also the basis of the apps and notebooks people who can't access DALL·E 2 are using.)