casual conversation
Open-Universe Assistance Games
Ma, Rachel, Qu, Jingyi, Bobu, Andreea, Hadfield-Menell, Dylan
Embodied AI agents must infer and act in an interpretable way on diverse human goals and preferences that are not predefined. To formalize this setting, we introduce Open-Universe Assistance Games (OU-AGs), a framework where the agent must reason over an unbounded and evolving space of possible goals. In this context, we introduce GOOD (GOals from Open-ended Dialogue), a data-efficient, online method that extracts goals in the form of natural language during an interaction with a human, and infers a distribution over natural language goals. GOOD prompts an LLM to simulate users with different complex intents, using its responses to perform probabilistic inference over candidate goals. This approach enables rich goal representations and uncertainty estimation without requiring large offline datasets. We evaluate GOOD in a text-based grocery shopping domain and in a text-operated simulated household robotics environment (AI2Thor), using synthetic user profiles. Our method outperforms a baseline without explicit goal tracking, as confirmed by both LLM-based and human evaluations.
The rise of AI companionship in a lonely Japan
Thirty-two and single, Akiho Sakai dreams of owning a cat to keep her company. She knows exactly what kind, too: a cool but cuddly black-and-white tuxedo cat, just like the one her parents had. The problem is, she can't. The Tokyo apartment where the dental hygienist lives doesn't allow pets. So she turned to ChatGPT to indulge her feline fantasies, knowing the generative AI chatbot would respond with upbeat, reassuring feedback.
MTP: A Dataset for Multi-Modal Turning Points in Casual Conversations
Ho, Gia-Bao Dinh, Tan, Chang Wei, Darban, Zahra Zamanzadeh, Salehi, Mahsa, Haffari, Gholamreza, Buntine, Wray
Detecting critical moments, such as emotional outbursts or changes in decisions during conversations, is crucial for understanding shifts in human behavior and their consequences. Our work introduces a novel problem setting focusing on these moments as turning points (TPs), accompanied by a meticulously curated, high-consensus, human-annotated multi-modal dataset. We provide precise timestamps, descriptions, and visual-textual evidence high-lighting changes in emotions, behaviors, perspectives, and decisions at these turning points. We also propose a framework, TPMaven, utilizing state-of-the-art vision-language models to construct a narrative from the videos and large language models to classify and detect turning points in our multi-modal dataset. Evaluation results show that TPMaven achieves an F1-score of 0.88 in classification and 0.61 in detection, with additional explanations aligning with human expectations.
Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
Ogawa, Atsunori, Kamo, Naoyuki, Matsuura, Kohei, Ashihara, Takanori, Moriya, Takafumi, Kano, Takatomo, Tawara, Naohiro, Delcroix, Marc
Large language models (LLMs) have been successfully applied for rescoring automatic speech recognition (ASR) hypotheses. However, their ability to rescore ASR hypotheses of casual conversations has not been sufficiently explored. In this study, we reveal it by performing N-best ASR hypotheses rescoring using Llama2 on the CHiME-7 distant ASR (DASR) task. Llama2 is one of the most representative LLMs, and the CHiME-7 DASR task provides datasets of casual conversations between multiple participants. We investigate the effects of domain adaptation of the LLM and context carry-over when performing N-best rescoring. Experimental results show that, even without domain adaptation, Llama2 outperforms a standard-size domain-adapted Transformer-LM, especially when using a long context. Domain adaptation shortens the context length needed with Llama2 to achieve its best performance, i.e., it reduces the computational cost of Llama2.
With a Little Help from my (Linguistic) Friends: Topic Segmentation of Multi-party Casual Conversations
Decker, Amandine, Amblard, Maxime
Topics play an important role in the global organisation of a conversation as what is currently discussed constrains the possible contributions of the participant. Understanding the way topics are organised in interaction would provide insight on the structure of dialogue beyond the sequence of utterances. However, studying this high-level structure is a complex task that we try to approach by first segmenting dialogues into smaller topically coherent sets of utterances. Understanding the interactions between these segments would then enable us to propose a model of topic organisation at a dialogue level. In this paper we work with open-domain conversations and try to reach a comparable level of accuracy as recent machine learning based topic segmentation models but with a formal approach. The features we identify as meaningful for this task help us understand better the topical structure of a conversation.
The Casual Conversations v2 Dataset
Porgali, Bilal, Albiero, Vรญtor, Ryda, Jordan, Ferrer, Cristian Canton, Hazirbas, Caner
This paper introduces a new large consent-driven dataset aimed at assisting in the evaluation of algorithmic bias and robustness of computer vision and audio speech models in regards to 11 attributes that are self-provided or labeled by trained annotators. The dataset includes 26,467 videos of 5,567 unique paid participants, with an average of almost 5 videos per person, recorded in Brazil, India, Indonesia, Mexico, Vietnam, Philippines, and the USA, representing diverse demographic characteristics. The participants agreed for their data to be used in assessing fairness of AI models and provided self-reported age, gender, language/dialect, disability status, physical adornments, physical attributes and geo-location information, while trained annotators labeled apparent skin tone using the Fitzpatrick Skin Type and Monk Skin Tone scales, and voice timbre. Annotators also labeled for different recording setups and per-second activity annotations.
ChatGPT: Why Everyone Is Obsessed This Mind-Blowing AI Chatbot โ Codelivly
There's a new chatbot in town, and it's causing quite a stir. ChatGPT is an artificial intelligence-powered chatbot that has garnered a lot of attention and hype in recent months. But what exactly is ChatGPT and why is everyone so obsessed with it? First and foremost, ChatGPT is a chatbot that utilizes the latest in artificial intelligence technology to converse with users in a natural and human-like manner. It can hold conversations on a wide range of topics, from current events to personal interests, and can even provide helpful recommendations or advice.
Facebook asked people to share their age and gender to create a fairer AI dataset
Facebook is sharing a new and diverse dataset with the wider AI community. In an announcement spotted by VentureBeat, the company says it envisions researchers using the collection, dubbed Casual Conversations, to test their machine learning models for bias. The dataset includes 3,011 people across 45,186 videos and gets its name from the fact it features those individuals providing unscripted answers to the company's questions. What's significant about Casual Conversations is that it involves paid actors who Facebook explicitly asked to share their age and gender. The company also hired trained professionals to label ambient lighting and the skin tones of those involved according to the Fitzpatrick scale, a dermatologist-developed system for classifying human skin colors.
Facebook could replace emoji with YOUMOJI
Emoticons allow Facebook years to truly express their feelings when words fail. But a new patent reveals the social media giant could be working on a method that shows your friends exactly how you feel. Facebook describes a computing device that detects an emoji in the text box and replaces it with the user's picture that corresponds with the meaning. Facebook describes a computing device that detects an emoji in the text box and replaces it with a user's picture that corresponds with the meaning, which could put an end to no more confusion about how you really feel This innovated patent was filed in March 2016 and was published a few months after โ July 28. These emoticons could be a combination of ':)' or ' 3' or even ':('- depending on how you are feeling at the time.
'Event[0]' Is a Fresh Take on the AI Story
The central hook to Event[0] is a character named Kaizen, a chatbot/AI who is surprisingly good at casual conversation. Other than you, the protagonist, Kaizen is literally the only other "living" thing in the game. There are other characters, sure, but they only exist in the past through chatlogs and recordings. Kaizen is all we have for companionship, so it's a damn good thing that the AI is such a compelling companion. You're originally part of a mission to Europa, but something goes wrong along the way, and you end up in an escape capsule, drifting through space. Eventually you drift your way to the Nautilus, a prototype ship now devoid of all life save for Kaizen, the caretaker AI.