Media
Time-aware Prompting for Text Generation
In this paper, we study the effects of incorporating timestamps, such as document creation dates, into generation systems. Two types of time-aware prompts are investigated: (1) textual prompts that encode document timestamps in natural language sentences; and (2) linear prompts that convert timestamps into continuous vectors. To explore extrapolation to future data points, we further introduce a new data-to-text generation dataset, TempWikiBio, containing more than 4 millions of chronologically ordered revisions of biographical articles from English Wikipedia, each paired with structured personal profiles. Through data-to-text generation on TempWikiBio, text-to-text generation on the content transfer dataset, and summarization on XSum, we show that linear prompts on encoder and textual prompts improve the generation quality on all datasets. Despite having less performance drop when testing on data drawn from a later time, linear prompts focus more on non-temporal information and are less sensitive to the given timestamps, according to human evaluations and sensitivity analyses. Meanwhile, textual prompts establish the association between the given timestamps and the output dates, yielding more factual temporal information in the output.
When to Laugh and How Hard? A Multimodal Approach to Detecting Humor and its Intensity
Alnajjar, Khalid, Hämäläinen, Mika, Tiedemann, Jörg, Laaksonen, Jorma, Kurimo, Mikko
Prerecorded laughter accompanying dialog in comedy TV shows encourages the audience to laugh by clearly marking humorous moments in the show. We present an approach for automatically detecting humor in the Friends TV show using multimodal data. Our model is capable of recognizing whether an utterance is humorous or not and assess the intensity of it. We use the prerecorded laughter in the show as annotation as it marks humor and the length of the audience's laughter tells us how funny a given joke is. We evaluate the model on episodes the model has not been exposed to during the training phase. Our results show that the model is capable of correctly detecting whether an utterance is humorous 78% of the time and how long the audience's laughter reaction should last with a mean absolute error of 600 milliseconds.
Speed Up the Cold-Start Learning in Two-Sided Bandits with Many Arms
Bayati, Mohsen, Cao, Junyu, Chen, Wanning
Multi-armed bandit (MAB) algorithms are efficient approaches to reduce the opportunity cost of online experimentation and are used by companies to find the best product from periodically refreshed product catalogs. However, these algorithms face the so-called cold-start at the onset of the experiment due to a lack of knowledge of customer preferences for new products, requiring an initial data collection phase known as the burn-in period. During this period, MAB algorithms operate like randomized experiments, incurring large burn-in costs which scale with the large number of products. We attempt to reduce the burn-in by identifying that many products can be cast into two-sided products, and then naturally model the rewards of the products with a matrix, whose rows and columns represent the two sides respectively. Next, we design two-phase bandit algorithms that first use subsampling and low-rank matrix estimation to obtain a substantially smaller targeted set of products and then apply a UCB procedure on the target products to find the best one. We theoretically show that the proposed algorithms lower costs and expedite the experiment in cases when there is limited experimentation time along with a large product set. Our analysis also reveals three regimes of long, short, and ultra-short horizon experiments, depending on dimensions of the matrix. Empirical evidence from both synthetic data and a real-world dataset on music streaming services validates this superior performance.
A Survey on Artificial Intelligence for Music Generation: Agents, Domains and Perspectives
Hernandez-Olivan, Carlos, Hernandez-Olivan, Javier, Beltran, Jose R.
Music is one of the Gardner's intelligences in his theory of multiple intelligences. How humans perceive and understand music is still being studied and is crucial to develop artificial intelligence models that imitate such processes. Music generation with Artificial Intelligence is an emerging field that is gaining much attention in the recent years. In this paper, we describe how humans compose music and how new AI systems could imitate such process by comparing past and recent advances in the field with music composition techniques. To understand how AI models and algorithms generate music and the potential applications that might appear in the future, we explore, analyze and describe the agents that take part of the music generation process: the datasets, models, interfaces, the users and the generated music. We mention possible applications that might benefit from this field and we also propose new trends and future research directions that could be explored in the future.
Comparing quantiles at scale in online A/B-testing - Spotify Engineering
TL;DR: Using the properties of the Poisson bootstrap algorithm and quantile estimators, we have been able to reduce the computational complexity of Poisson bootstrap difference-in-quantiles confidence intervals enough to unlock bootstrap inference for almost arbitrary large samples. At Spotify, we can now easily calculate bootstrap confidence intervals for difference-in-quantiles in A/B tests with hundreds of millions of observations. In product development, the most common impact analysis of product changes is often summarized by the change in the average of some metric of interest. This is a natural measurement, since changes in an average, in many contexts, map more or less directly to changes in business value. In addition, averages have convenient mathematical properties that make it straightforward to quantify uncertainty in over-served changes.
Make it pop! Do we really need the Beatles to sound new?
Yellow Submarine, Ringo Starr's turn on Revolver, has been a gateway for children into the music of the Beatles since its release in 1966. A new reissue of the album makes that relationship more explicit: Giles Martin, son of original producer George and the sonic custodian of the Beatles catalogue, says his "de-mixing" of the album – using AI to separate individual instruments that were originally squeezed together on four tracks – was done in part with a playlist-listening younger audience in mind. Martin recently told Variety that his teenage children listen to old and new music side by side, veering from Fleetwood Mac to Billie Eilish and Olivia Rodrigo. "[W]hat I want to make sure is that when people hear the Beatles, that it has the same dynamic as the other stuff they're listening to," he said. He added that 1969's Abbey Road, recorded on a then luxuriant eight tracks and the first Beatles album not released in mono, stands out from the band's catalogue as "it sounds more hi-fi than the other Beatles albums".
AI analysis of segments on CNN, Fox News and MSNBC shows females get less airtime
Artificial intelligence has found disparities in the amount of airtime women and men were given on CNN, FOX News and MSNBC - females had a 10 percent less chance of speaking during political discussions because male speakers constantly interrupted them. The discovery was made by researchers at Rochester Institute of Technology who analyzed 625,409 dialogues hosted on the three news cable networks from January 2000 through July 2021. The technology revealed women received an average of 72.8 words per chance to speak compared to 81.4 for male speakers and women were interrupted 39.4 percent of the time during discussions - this is compared to the 35.9 percent of the time for men. The team believes their AI could be used during talk shows, interviews and political debates to identify a serial interrupter in real-time, but the study also reinforces previous research that found men interrupt women more to show their dominance. AI analyzed thousands of dialogues from news segments on the three networks and found woman are given a 10 percent less chance at speaking because men interrupt them.
The X-T5 is the first major upgrade to Fujifilm's compact camera flagship in 5 years
Fujifilm is delivering a follow-up to the well-received X-T4. The company has introduced (what else?) the X-T5, a sequel to the higher-end APS-C mirrorless camera that delivers some major technical upgrades -- the largest in five years -- while refining the basic formula. The new model now packs Fuji's current 40MP sensor (up from 26MP) that can shoot 6.2K video at 30 frames per second. You don't need to buy a top-tier cam like the X-H2S to venture beyond 4K. You can also expect a jump in computing power through the X-Processor 5 that allows for AI-based autofocusing, 4:2:2 10-bit output, F-log2 and support for the HEIF photo format.