Media
Simulating Social Media Using Large Language Models to Evaluate Alternative News Feed Algorithms
Törnberg, Petter, Valeeva, Diliara, Uitermark, Justus, Bail, Christopher
Social media is often criticized for amplifying toxic discourse and discouraging constructive conversations. But designing social media platforms to promote better conversations is inherently challenging. This paper asks whether simulating social media through a combination of Large Language Models (LLM) and Agent-Based Modeling can help researchers study how different news feed algorithms shape the quality of online conversations. We create realistic personas using data from the American National Election Study to populate simulated social media platforms. Next, we prompt the agents to read and share news articles - and like or comment upon each other's messages - within three platforms that use different news feed algorithms. In the first platform, users see the most liked and commented posts from users whom they follow. In the second, they see posts from all users - even those outside their own network. The third platform employs a novel "bridging" algorithm that highlights posts that are liked by people with opposing political views. We find this bridging algorithm promotes more constructive, non-toxic, conversation across political divides than the other two models. Though further research is needed to evaluate these findings, we argue that LLMs hold considerable potential to improve simulation research on social media and many other complex social settings.
CineTransfer: Controlling a Robot to Imitate Cinematographic Style from a Single Example
Pueyo, Pablo, Montijano, Eduardo, Murillo, Ana C., Schwager, Mac
This work presents CineTransfer, an algorithmic framework that drives a robot to record a video sequence that mimics the cinematographic style of an input video. We propose features that abstract the aesthetic style of the input video, so the robot can transfer this style to a scene with visual details that are significantly different from the input video. The framework builds upon CineMPC, a tool that allows users to control cinematographic features, like subjects' position on the image and the depth of field, by manipulating the intrinsics and extrinsics of a cinematographic camera. However, CineMPC requires a human expert to specify the desired style of the shot (composition, camera motion, zoom, focus, etc). CineTransfer bridges this gap, aiming a fully autonomous cinematographic platform. The user chooses a single input video as a style guide. CineTransfer extracts and optimizes two important style features, the composition of the subject in the image and the scene depth of field, and provides instructions for CineMPC to control the robot to record an output sequence that matches these features as closely as possible. In contrast with other style transfer methods, our approach is a lightweight and portable framework which does not require deep network training or extensive datasets. Experiments with real and simulated videos demonstrate the system's ability to analyze and transfer style between recordings, and are available in the supplementary video.
Deep Generative Models of Music Expectation
Masclef, Ninon Lizé, Keller, T. Anderson
A prominent theory of affective response to music revolves around the concepts of surprisal and expectation. In prior work, this idea has been operationalized in the form of probabilistic models of music which allow for precise computation of song (or note-by-note) probabilities, conditioned on a 'training set' of prior musical or cultural experiences. To date, however, these models have been limited to compute exact probabilities through hand-crafted features or restricted to linear models which are likely not sufficient to represent the complex conditional distributions present in music. In this work, we propose to use modern deep probabilistic generative models in the form of a Diffusion Model to compute an approximate likelihood of a musical input sequence. Unlike prior work, such a generative model parameterized by deep neural networks is able to learn complex non-linear features directly from a training set itself. In doing so, we expect to find that such models are able to more accurately represent the 'surprisal' of music for human listeners. From the literature, it is known that there is an inverted U-shaped relationship between surprisal and the amount human subjects 'like' a given song. In this work we show that pre-trained diffusion models indeed yield musical surprisal values which exhibit a negative quadratic relationship with measured subject 'liking' ratings, and that the quality of this relationship is competitive with state of the art methods such as IDyOM. We therefore present this model a preliminary step in developing modern deep generative models of music expectation and subjective likability.
USB-NeRF: Unrolling Shutter Bundle Adjusted Neural Radiance Fields
Li, Moyang, Wang, Peng, Zhao, Lingzhe, Liao, Bangyan, Liu, Peidong
Neural Radiance Fields (NeRF) has received much attention recently due to its impressive capability to represent 3D scene and synthesize novel view images. Existing works usually assume that the input images are captured by a global shutter camera. Thus, rolling shutter (RS) images cannot be trivially applied to an off-the-shelf NeRF algorithm for novel view synthesis. Rolling shutter effect would also affect the accuracy of the camera pose estimation (e.g. via COLMAP), which further prevents the success of NeRF algorithm with RS images. In this paper, we propose Unrolling Shutter Bundle Adjusted Neural Radiance Fields (USB-NeRF). USB-NeRF is able to correct rolling shutter distortions and recover accurate camera motion trajectory simultaneously under the framework of NeRF, by modeling the physical image formation process of a RS camera. Experimental results demonstrate that USB-NeRF achieves better performance compared to prior works, in terms of RS effect removal, novel view image synthesis as well as camera motion estimation. Furthermore, our algorithm can also be used to recover high-fidelity high frame-rate global shutter video from a sequence of RS images. Understanding and recovering 3D scenes from 2D images is a difficult but important problem in computer vision. Different from a 2D image which can be naturally formulated as an array of pixel values, there are many 3D representations to depict a 3D scene, such as the commonly used point clouds (Furukawa & Ponce, 2009), height-map (Pollefeys et al., 2008), voxel grids (Nießner et al., 2013; Seitz & Dyer, 1997) and 3D triangular meshes (Delaunoy & Pollefeys, 2014).
DISCO-10M: A Large-Scale Music Dataset
Lanzendörfer, Luca A., Grötschla, Florian, Funke, Emil, Wattenhofer, Roger
Music datasets play a crucial role in advancing research in machine learning for music. However, existing music datasets suffer from limited size, accessibility, and lack of audio resources. To address these shortcomings, we present DISCO-10M, a novel and extensive music dataset that surpasses the largest previously available music dataset by an order of magnitude. To ensure high-quality data, we implement a multi-stage filtering process. This process incorporates similarities based on textual descriptions and audio embeddings. Moreover, we provide precomputed CLAP embeddings alongside DISCO-10M, facilitating direct application on various downstream tasks. These embeddings enable efficient exploration of machine learning applications on the provided data. With DISCO-10M, we aim to democratize and facilitate new research to help advance the development of novel machine learning models for music.
In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT
Shen, Xinyue, Chen, Zeyuan, Backes, Michael, Zhang, Yang
The way users acquire information is undergoing a paradigm shift with the advent of ChatGPT. Unlike conventional search engines, ChatGPT retrieves knowledge from the model itself and generates answers for users. ChatGPT's impressive question-answering (QA) capability has attracted more than 100 million users within a short period of time but has also raised concerns regarding its reliability. In this paper, we perform the first large-scale measurement of ChatGPT's reliability in the generic QA scenario with a carefully curated set of 5,695 questions across ten datasets and eight domains. We find that ChatGPT's reliability varies across different domains, especially underperforming in law and science questions. We also demonstrate that system roles, originally designed by OpenAI to allow users to steer ChatGPT's behavior, can impact ChatGPT's reliability in an imperceptible way. We further show that ChatGPT is vulnerable to adversarial examples, and even a single character change can negatively affect its reliability in certain cases. We believe that our study provides valuable insights into ChatGPT's reliability and underscores the need for strengthening the reliability and security of large language models (LLMs).
Meta introduces generative AI tools for advertisers to enhance content creation
Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. Social media giant Meta Platforms said on Wednesday that it has started rolling out generative artificial intelligence (AI) tools that can create content like image backgrounds and variations of written text for all advertisers. The company started testing these tools in May, giving access to a select group of advertisers in a "testing playground". The tools will be available in Meta's Ads Manager and their rollout will be completed next year.
How to use AI to help you get a better job instead of it stealing one
CyberGuy shows you how to manage your online presence. The job search landscape has transformed dramatically in just a few years. Gone are the days when applying for jobs was a part-time endeavor. Nowadays, it's practically a full-time job, especially if you're out of work and have to document your efforts to claim unemployment benefits. The experience can be overwhelming, but fortunately, technology--particularly artificial intelligence (AI)--is here to help streamline the process.
Google Pixel 8 gets more nifty AI-powered editing tools for photo and video
Google's hardware event has been chock full of information on new devices, like the Pixel 8 smartphone, but camera software has also gotten some TLC. The company announced a ton of Pixel 8 features exclusive for shutterbugs and video editors. The new Best Take feature solves the issue of, uh, one person looking really gross in group photos. When enabled, the software takes a series of photos in quick succession and you can actually mix and match faces to create the perfect group shot, sort of a face-based riff on the pre-existing Magic Editor tech. Grab a face from one photo and slap it on the next.
The Origin Story of "Stop Making Sense"
When it first opened in theatres, in the fall of 1984, "Stop Making Sense," directed by Jonathan Demme and starring the rock group Talking Heads, was quickly recognized as one of the finest concert films ever made. Reviewer after reviewer settled on the word "exhilarating" to describe the experience of watching an expanded nine-member iteration of the four-piece group perform sixteen of their best-known songs in an uninterrupted sequence of dynamically staged and photographed musical vignettes. In the pages of this magazine, Pauline Kael praised the film as "close to perfection," and described the Heads front man, David Byrne, as "a stupefying performer." "He's so white he's almost mock-white," Kael wrote, "and so are his jerky, long-necked, mechanical-man movements. He seems fleshless, bloodless; he might almost be a Black man's parody of how a clean-cut white man moves. But Byrne himself is the parodist, and he commands the stage by his hollow-eyed, frosty verve."