Goto

Collaborating Authors

 Media


What the AI Chatbot Discourse Is Really Revealing

#artificialintelligence

The biggest tech story of the year is shaping up around the seemingly sudden arrival of AI chatbots into mainstream attention: piggybacking off last year's viral reception to text-to-image generators like DALL-E 2, the launch of OpenAI's ChatGPT in November has since spurred not only widespread media coverage and netizen adoption, but also an industry-wide arms race. Whatever polite corporate doffing made to AI's thicket of ethical ramifications over the past few decades disintegrated nearly overnight in favor of Silicon Valley's primal fear of competition, and we now live in a society where Microsoft's newly AI-powered Bing ("Sydney," to her friends), Google's Bard, Meta's LLaMA, and Snapchat's My AI (which at least allows you the dignity of naming your chatbot yourself) seem poised to transform us all. The AI future feels nigh, if not terribly optimistic. In an era where major breakthroughs in tech render either inscrutable--admit it, you still don't know what a blockchain is, do you?--or We're kind of used to it already: After spending the greater part of Web 2.0 accepting the sleight of hand that invisible, algorithmic forces exert on our day-to-day, the consumer-friendly AI-powered machinations of driverless cars and actually efficient task assistants and decent predictive-text features has become a foregone conclusion.


Welcome to the Museum of the Future AI Apocalypse

WIRED

Audrey Kim's dog Murphy uses a combination of head nods and 10 buttons on the ground to communicate, she says, and has a habit of making friends with crows. She taught him to use the buttons because she believes consciousness is a spectrum and intelligence is mysterious. Those tenets also led her to become curator of the Misalignment Museum, a temporary exhibition about the future of artificial intelligence that opens today in San Francisco, ground zero for recent excitement about generative AI and chatbots like OpenAI's ChatGPT. The Misalignment Museum imagines a future in which AI starts to take the route mapped out in countless science fiction films--becoming self aware and setting about killing off humanity. Fortunately, in Kim's vision the algorithms self-correct and stop short of killing all people.


From Ghostbusters and Aliens to Lego Star Wars: 10 great video games based on movies

The Guardian

Designed by ex-Atari luminary David Crane (Pitfall, Decathlon), Activision's wonderful tie-in captured the humour and spirit of the classic comedy. Players set up their own ghostbusting franchises, buying equipment before setting out to capture spooks. With its use of digitised speech and a jaunty reproduction of the film's soundtrack, it showed that games really could provide an authentic movie experience. Developed by the UK-based movie tie-in specialist Probe Entertainment, Die Hard Trilogy is three games in one: a third-person action adventure, a light-gun shooter and an arcade driving challenge, each based around consecutive instalments of the film series. Though the visuals were rough, the game perfectly captured the locations, themes and black humour of the movies, providing a real bargain for early PlayStation and Saturn owners.


Co-Speech Gesture Synthesis using Discrete Gesture Token Learning

arXiv.org Artificial Intelligence

Synthesizing realistic co-speech gestures is an important and yet unsolved problem for creating believable motions that can drive a humanoid robot to interact and communicate with human users. Such capability will improve the impressions of the robots by human users and will find applications in education, training, and medical services. One challenge in learning the co-speech gesture model is that there may be multiple viable gesture motions for the same speech utterance. The deterministic regression methods can not resolve the conflicting samples and may produce over-smoothed or damped motions. We proposed a two-stage model to address this uncertainty issue in gesture synthesis by modeling the gesture segments as discrete latent codes. Our method utilizes RQ-VAE in the first stage to learn a discrete codebook consisting of gesture tokens from training data. In the second stage, a two-level autoregressive transformer model is used to learn the prior distribution of residual codes conditioned on input speech context. Since the inference is formulated as token sampling, multiple gesture sequences could be generated given the same speech input using top-k sampling. The quantitative results and the user study showed the proposed method outperforms the previous methods and is able to generate realistic and diverse gesture motions.


Personalized Reward Learning with Interaction-Grounded Learning (IGL)

arXiv.org Artificial Intelligence

In an era of countless content offerings, recommender systems alleviate information overload by providing users with personalized content suggestions. Due to the scarcity of explicit user feedback, modern recommender systems typically optimize for the same fixed combination of implicit feedback signals across all users. However, this approach disregards a growing body of work highlighting that (i) implicit signals can be used by users in diverse ways, signaling anything from satisfaction to active dislike, and (ii) different users communicate preferences in different ways. We propose applying the recent Interaction Grounded Learning (IGL) paradigm to address the challenge of learning representations of diverse user communication modalities. Rather than requiring a fixed, human-designed reward function, IGL is able to learn personalized reward functions for different users and then optimize directly for the latent user satisfaction. We demonstrate the success of IGL with experiments using simulations as well as with real-world production traces. From shopping to reading the news, modern Internet users have access to an overwhelming amount of content and choices from online services. Recommender systems offer a way to improve user experience and decrease information overload by providing a customized selection of content. A key challenge for recommender systems is the rarity of explicit user feedback, such as ratings or likes/dislikes (Grฤar et al., 2005). Rather than explicit feedback, practitioners typically use more readily available implicit signals, such as clicks (Hu et al., 2008), webpage dwell time (Yi et al., 2014), or inter-arrival times (Wu et al., 2017) as a proxy signal for user satisfaction. These implicit signals are used as the reward objective in recommender systems, with the popular Click-Through Rate (CTR) metric as the gold standard for the field (Silveira et al., 2019).


Who could be behind QAnon? Authorship attribution with supervised machine-learning

arXiv.org Artificial Intelligence

A series of social media posts signed under the pseudonym "Q", started a movement known as QAnon, which led some of its most radical supporters to violent and illegal actions. To identify the person(s) behind Q, we evaluate the coincidence between the linguistic properties of the texts written by Q and to those written by a list of suspects provided by journalistic investigation. To identify the authors of these posts, serious challenges have to be addressed. The "Q drops" are very short texts, written in a way that constitute a sort of literary genre in itself, with very peculiar features of style. These texts might have been written by different authors, whose other writings are often hard to find. After an online ethnology of the movement, necessary to collect enough material written by these thirteen potential authors, we use supervised machine learning to build stylistic profiles for each of them. We then performed a rolling analysis on Q's writings, to see if any of those linguistic profiles match the so-called 'QDrops' in part or entirety. We conclude that two different individuals, Paul F. and Ron W., are the closest match to Q's linguistic signature, and they could have successively written Q's texts. These potential authors are not high-ranked personality from the U.S. administration, but rather social media activists.


anafi_ros: from Off-the-Shelf Drones to Research Platforms

arXiv.org Artificial Intelligence

The off-the-shelf drones are simple to operate and easy to maintain aerial systems. However, due to proprietary flight software, these drones usually do not provide any open-source interface which can enable them for autonomous flight in research or teaching. This work introduces a package for ROS1 and ROS2 for straightforward interfacing with off-the-shelf drones from the Parrot ANAFI family. The developed ROS package is hardware agnostic, allowing connecting seamlessly to all four supported drone models. This framework can connect with the same ease to a single drone or a team of drones from the same ground station. The developed package was intensively tested at the limits of the drones' capabilities and thoughtfully documented to facilitate its use by other research groups worldwide.


CONTAIN: A Community-based Algorithm for Network Immunization

arXiv.org Artificial Intelligence

The adoption of advanced digital technologies has transformed and evolved social media, which in turn enhanced the connectivity and awareness of our society. Along with these advancements, digitalization also created a favorable environment for the diffusion of misinformation [9, 22, 27, 26] (i.e., the unintentional spread of false information), disinformation (i.e., the intentional spread of false information), and hate speech (i.e., the intentional spread of malicious content expressing hate and violence). Conversely, research in network analysis has advanced, proposing various ideas for network immunization strategies [30, 4, 15, 31, 16]. Given a social network, consider how the content produced from a node spreads in the network. This is known as information diffusion, and it has shown to be a crucial field of research in network analysis - as an example, how the political ideas of a politician influence a social network [1].


CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos

arXiv.org Artificial Intelligence

Recent years have seen progress beyond domain-specific sound separation for speech or music towards universal sound separation for arbitrary sounds. Prior work on universal sound separation has investigated separating a target sound out of an audio mixture given a text query. Such text-queried sound separation systems provide a natural and scalable interface for specifying arbitrary target sounds. However, supervised text-queried sound separation systems require costly labeled audio-text pairs for training. Moreover, the audio provided in existing datasets is often recorded in a controlled environment, causing a considerable generalization gap to noisy audio in the wild. In this work, we aim to approach text-queried universal sound separation by using only unlabeled data. We propose to leverage the visual modality as a bridge to learn the desired audio-textual correspondence. The proposed CLIPSep model first encodes the input query into a query vector using the contrastive language-image pretraining (CLIP) model, and the query vector is then used to condition an audio separation model to separate out the target sound. While the model is trained on image-audio pairs extracted from unlabeled videos, at test time we can instead query the model with text inputs in a zero-shot setting, thanks to the joint language-image embedding learned by the CLIP model. Further, videos in the wild often contain off-screen sounds and background noise that may hinder the model from learning the desired audio-textual correspondence. To address this problem, we further propose an approach called noise invariant training for training a query-based sound separation model on noisy data. Experimental results show that the proposed models successfully learn text-queried universal sound separation using only noisy unlabeled videos, even achieving competitive performance against a supervised model in some settings.


Meme Sentiment Analysis Enhanced with Multimodal Spatial Encoding and Facial Embedding

arXiv.org Artificial Intelligence

Internet memes are characterised by the interspersing of text amongst visual elements. State-of-the-art multimodal meme classifiers do not account for the relative positions of these elements across the two modalities, despite the latent meaning associated with where text and visual elements are placed. Against two meme sentiment classification datasets, we systematically show performance gains from incorporating the spatial position of visual objects, faces, and text clusters extracted from memes. In addition, we also present facial embedding as an impactful enhancement to image representation in a multimodal meme classifier. Finally, we show that incorporating this spatial information allows our fully automated approaches to outperform their corresponding baselines that rely on additional human validation of OCR-extracted text.