Media
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Zhang, Yue, Li, Yafu, Cui, Leyang, Cai, Deng, Liu, Lemao, Fu, Tingchen, Huang, Xinting, Zhao, Enbo, Zhang, Yu, Chen, Yulong, Wang, Longyue, Luu, Anh Tuan, Bi, Wei, Shi, Freda, Shi, Shuming
While large language models (LLMs) have demonstrated remarkable capabilities across a range of downstream tasks, a significant concern revolves around their propensity to exhibit hallucinations: LLMs occasionally generate content that diverges from the user input, contradicts previously generated context, or misaligns with established world knowledge. This phenomenon poses a substantial challenge to the reliability of LLMs in real-world scenarios. In this paper, we survey recent efforts on the detection, explanation, and mitigation of hallucination, with an emphasis on the unique challenges posed by LLMs. We present taxonomies of the LLM hallucination phenomena and evaluation benchmarks, analyze existing approaches aiming at mitigating LLM hallucination, and discuss potential directions for future research.
Beyond Deep Fakes
Within the next five years, the way we work, live, play, and learn will be changed by digital humans (chatbots and avatars with very realistic human faces). Digital humans are already gaining popularity as social media influencers, and they will soon evolve into digital sales assistants, fashion advisers, and personal shoppers able to model how customers will look and move in the latest ensembles. Digital humans will become central to the multibillion-dollar fashion industry, as social media is further integrated into the retail customer experience. Digital humans will also help in healthcare, enabling medical students and social workers to develop better interview skills for patients in sensitive clinical settings. They will allow people, especially those with mental health challenges, to rehearse for job interviews. They will help keep elderly people connected to their communities and respectfully monitored so they can remain in their homes longer. They will provide a human face for personalized advice, support, and training--and do it at scale. This has become possible with the advent of cost-effective, highly realistic, personalized interactive digital agents and avatars sporting high-fidelity facial simulations powered by advances in both real-time neural rendering (NR) and low-latency computing. NR refers to the use of machine-learning (ML) techniques to generate digital faces or face replacements in video.17 NR rose to prominence with the advent of so-called "deep fakes"--the replacement of someone's face in videos with an NR-generated face of remarkable realism. The term originates from the name of a Reddit user (/u/deepfakes), a ML engineer who posted the original deep fake auto-encoder. Often used for satire, deep fakes can be harmful, presenting novel ethical issues. The best-known examples involve deep fakes of celebrities, a form of face "hijacking" whereby publicly available videos of a person are used to train an ML program that overlays the source person's face onto existing video footage; this technique was originally used in pornographic material.
An AI Game of Thrones prequel? No wonder George RR Martin's raining ice and fire on ChatGPT Tim Adams
Battles between human and artificial intelligence are no longer science fiction. The strikes in Hollywood led by the united guilds of actors and screenwriters have a common, intangible enemy: the algorithms and computer-generated imagery that are increasingly programmed by studios to render them redundant. In New York last week, a new front in that stand-off was opened by a group of American novelists โ including John Grisham, Jodi Picoult and Jonathan Franzen โ who are suing OpenAI, the creators of the ChatGPT program. The legal case may help to define and protect those increasingly porous boundaries between human creativity and the robots that mimic it. In the meantime, Amazon, these days flooded by self-published books written by AI, has taken its first half-hearted steps to curtail that practice.
What is Your Face Worth?
Felix Salmon, Emily Peck, and Elizabeth Spiers are joined by Kashmir Hill to talk about her new book, Your Face Belongs to Us. They dig into the way facial recognition technology is used in unexpected (and sometimes creepy) ways. They also talk about the A.I. revolution and Rupert Murdoch's "exit" from the Fox empire. If you enjoy this show, please consider signing up for Slate Plus. Slate Plus members get an ad-free experience across the network and an additional segment of our show every week.
Cordyceps@LT-EDI: Depression Detection with Reddit and Self-training
Depression is debilitating, and not uncommon. Indeed, studies of excessive social media users show correlations with depression, ADHD, and other mental health concerns. Given that there is a large number of people with excessive social media usage, then there is a significant population of potentially undiagnosed users and posts that they create. In this paper, we propose a depression severity detection system using a semi-supervised learning technique to predict if a post is from a user who is experiencing severe, moderate, or low (non-diagnostic) levels of depression. Namely, we use a trained model to classify a large number of unlabelled social media posts from Reddit, then use these generated labels to train a more powerful classifier. We demonstrate our framework on Detecting Signs of Depression from Social Media Text - LT-EDI@RANLP 2023 shared task, where our framework ranks 3rd overall.
From Text to Source: Results in Detecting Large Language Model-Generated Content
Antoun, Wissam, Sagot, Benoรฎt, Seddah, Djamรฉ
The widespread use of Large Language Models (LLMs), celebrated for their ability to generate human-like text, has raised concerns about misinformation and ethical implications. Addressing these concerns necessitates the development of robust methods to detect and attribute text generated by LLMs. This paper investigates "Cross-Model Detection," evaluating whether a classifier trained to distinguish between source LLM-generated and human-written text can also detect text from a target LLM without further training. The study comprehensively explores various LLM sizes and families, and assesses the impact of conversational fine-tuning techniques on classifier generalization. The research also delves into Model Attribution, encompassing source model identification, model family classification, and model size classification. Our results reveal several key findings: a clear inverse relationship between classifier effectiveness and model size, with larger LLMs being more challenging to detect, especially when the classifier is trained on data from smaller models. Training on data from similarly sized LLMs can improve detection performance from larger models but may lead to decreased performance when dealing with smaller models. Additionally, model attribution experiments show promising results in identifying source models and model families, highlighting detectable signatures in LLM-generated text. Overall, our study contributes valuable insights into the interplay of model size, family, and training data in LLM detection and attribution.
WikiMT++ Dataset Card
Zhou, Monan, Wu, Shangda, Wang, Yuan, Li, Wei
Table 1 shows the specific names and number of classes of genre and emotion labels. WikiMT++ is an expanded and refined version of WikiMusicText (WikiMT), featuring 1010 curated lead 2.1 Attributes from WikiMT or Information sheets in ABC notation. To expand application scenarios of The titles, artists, genres, and descriptions are directly inherited WikiMT, we add both objective (album, lyrics, video) and from WikiMT. However, as they were originally subjective emotion (12 emotion adjectives) and emo_4q curated from openly accessible sources, potential constraints (Russell 4Q) attributes, enhancing its usability for music and wrongs still exist. For better precision and information retrieval, conditional music generation, automatic completeness, we update these attributes through CLaMP composition, and emotion classification, etc.
Language-Guided Audio-Visual Source Separation via Trimodal Consistency
Tan, Reuben, Ray, Arijit, Burns, Andrea, Plummer, Bryan A., Salamon, Justin, Nieto, Oriol, Russell, Bryan, Saenko, Kate
We propose a self-supervised approach for learning to perform audio source separation in videos based on natural language queries, using only unlabeled video and audio pairs as training data. A key challenge in this task is learning to associate the linguistic description of a sound-emitting object to its visual features and the corresponding components of the audio waveform, all without access to annotations during training. To overcome this challenge, we adapt off-the-shelf vision-language foundation models to provide pseudo-target supervision via two novel loss functions and encourage a stronger alignment between the audio, visual and natural language modalities. During inference, our approach can separate sounds given text, video and audio input, or given text and audio input alone. We demonstrate the effectiveness of our self-supervised approach on three audio-visual separation datasets, including MUSIC, SOLOS and AudioSet, where we outperform state-of-the-art strongly supervised approaches despite not using object detectors or text labels during training.
Spotify's priciest lossless audio plan could sort playlists by "danceability"
Since 2017, there have been endless rumors and even a rescinded announcement of a HiFi tier at Spotify, but no option itself. A Reddit user has dug into Spotify's app and uncovered possible information about a Supremium tier (Spotify's new name for the HiFi option). Apparently, it could have 24-bit Lossless music, which the company claims is free from "lag and delays." Despite its uncertainty, some pretty fun features are currently floating around in that code, including the ability to sort playlists by "danceability." The option to determine how much you want to boogie could come alongside other arrangements like BPM and smart order, which would attempt to create an ideal playlist based on tempo and key.