Media
Assessing the Reasoning Abilities of ChatGPT in the Context of Claim Verification
Dougrez-Lewis, John, Akhter, Mahmud Elahi, He, Yulan, Liakata, Maria
The reasoning capabilities of LLMs are currently hotly debated. We examine the issue from the perspective of claim/rumour verification. We propose the first logical reasoning framework designed to break down any claim or rumor paired with evidence into the atomic reasoning steps necessary for verification. Based on our framework, we curate two annotated collections of such claim/evidence pairs: a synthetic dataset from Wikipedia and a real-world set stemming from rumours circulating on Twitter. We use them to evaluate the reasoning capabilities of GPT-3.5-Turbo and GPT-4 (hereinafter referred to as ChatGPT) within the context of our framework, providing a thorough analysis. Our results show that ChatGPT struggles in abductive reasoning, although this can be somewhat mitigated by using manual Chain of Thought (CoT) as opposed to Zero Shot (ZS) and ZS CoT approaches. Our study contributes to the growing body of research suggesting that ChatGPT's reasoning processes are unlikely to mirror human-like reasoning, and that LLMs need to be more rigorously evaluated in order to distinguish between hype and actual capabilities, especially in high stake real-world tasks such as claim verification.
Retrieve Only When It Needs: Adaptive Retrieval Augmentation for Hallucination Mitigation in Large Language Models
Ding, Hanxing, Pang, Liang, Wei, Zihao, Shen, Huawei, Cheng, Xueqi
Hallucinations pose a significant challenge for the practical implementation of large language models (LLMs). The utilization of parametric knowledge in generating factual content is constrained by the limited knowledge of LLMs, potentially resulting in internal hallucinations. While incorporating external information can help fill knowledge gaps, it also introduces the risk of irrelevant information, thereby increasing the likelihood of external hallucinations. A careful and balanced integration of the parametric knowledge within LLMs with external information is crucial to alleviate hallucinations. In this study, we present Rowen, a novel approach that enhances LLMs with a selective retrieval augmentation process tailored to address hallucinated outputs. This process is governed by a multilingual semantic-aware detection module, which evaluates the consistency of the perturbed responses across various languages for the same queries. Upon detecting inconsistencies indicative of hallucinations, Rowen activates the retrieval of external information to rectify the model outputs. Rowen adeptly harmonizes the intrinsic parameters in LLMs with external knowledge sources, effectively mitigating hallucinations by ensuring a balanced integration of internal reasoning and external evidence. Through a comprehensive empirical analysis, we demonstrate that Rowen surpasses the current state-of-the-art in both detecting and mitigating hallucinated content within the outputs of LLMs.
Can We Verify Step by Step for Incorrect Answer Detection?
Xu, Xin, Diao, Shizhe, Yang, Can, Wang, Yang
Chain-of-Thought (CoT) prompting has marked a significant advancement in enhancing the reasoning capabilities of large language models (LLMs). Previous studies have developed various extensions of CoT, which focus primarily on enhancing end-task performance. In addition, there has been research on assessing the quality of reasoning chains in CoT. This raises an intriguing question: Is it possible to predict the accuracy of LLM outputs by scrutinizing the reasoning chains they generate? To answer this research question, we introduce a benchmark, R2PE, designed specifically to explore the relationship between reasoning chains and performance in various reasoning tasks spanning five different domains. This benchmark aims to measure the falsehood of the final output of LLMs based on the reasoning steps. To make full use of information in multiple reasoning chains, we propose the process discernibility score (PDS) framework that beats the answer-checking baseline by a large margin. Concretely, this resulted in an average of 5.1% increase in the F1 score across all 45 subsets within R2PE. We further demonstrate our PDS's efficacy in advancing open-domain QA accuracy. Data and code are available at https://github.com/XinXU-USTC/R2PE.
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators
Cao, Yang Trista, Domingo, Lovely-Frances, Gilbert, Sarah Ann, Mazurek, Michelle, Shilton, Katie, Daumé, Hal III
Extensive efforts in automated approaches for content moderation have been focused on developing models to identify toxic, offensive, and hateful content with the aim of lightening the load for moderators. Yet, it remains uncertain whether improvements on those tasks have truly addressed moderators' needs in accomplishing their work. In this paper, we surface gaps between past research efforts that have aimed to provide automation for aspects of content moderation and the needs of volunteer content moderators, regarding identifying violations of various moderation rules. To do so, we conduct a model review on Hugging Face to reveal the availability of models to cover various moderation rules and guidelines from three exemplar forums. We further put state-of-the-art LLMs to the test, evaluating how well these models perform in flagging violations of platform rules from one particular forum. Finally, we conduct a user survey study with volunteer moderators to gain insight into their perspectives on useful moderation models. Overall, we observe a non-trivial gap, as missing developed models and LLMs exhibit moderate to low performance on a significant portion of the rules. Moderators' reports provide guides for future work on developing moderation assistant models.
The AI Industry Is Stuck on One Very Specific Way to Use a Chatbot
A perfect day in Los Angeles starts with a stroll along the Venice Beach boardwalk. After that, Beverly Hills, then Hollywood to see the Walk of Fame, then Griffith Park for a hike, then Chinatown for dim sum, then downtown, perhaps to catch an evening show at the Walt Disney Concert Hall. Or at least, that's what a chatbot thinks a "perfect day" is. This agenda was custom-made for me by Microsoft Copilot after I told it I had one day in town to explore the sights and asked it to plan accordingly. Here's a jam-packed 24-hour itinerary," Copilot responded, before rattling off an eight-part answer. What I didn't tell Copilot is that I already live here--and know that such an itinerary is perfect only if your idea of bliss is spending most of the day traversing one of the country's most sprawling, traffic-clogged cities, frantically popping from landmark to landmark. I asked Copilot to make me a travel itinerary because Microsoft has trotted it out as an example of how people can use the ChatGPT-like assistant. It can supposedly help you pick a destination, compare flight prices, and settle on attractions that are "popular with tourists--or just a little more off the beaten path." Of all the things you might ask a chatbot, AI companies love to suggest you ask for help planning upcoming travel. Open up ChatGPT and you might see this hypothetical prompt: "Plan a trip to see the best of New York in 3 days." Google's Gemini chatbot offers similar ones. Meta's line of chatbot assistants on Instagram and Facebook includes "Lorena," your own personal travel expert. And Rabbit, the company behind a new AI gadget, pulled out the travel example for a keynote video last month. If one were to play AI-marketing bingo, "trip itinerary" would get crossed off basically every time. More than a year into the generative-AI revolution, companies so frequently suggest that people use their tools in this way that you'd think chatbots would excel at it. In theory, chatbots that can instantaneously create travel plans are a marketer's dream. The use case is easy to understand: Planning a vacation can be a real challenge for people. First, it involves toggling among flight listings, hotel availability, and ticketing websites for major attractions. Then, it requires more nuanced research, to figure out which local restaurants are actually good and which are overpriced tourist scams, or what time to set off for a big hike that won't leave you in the woods after sunset. Most of this travel information already lives on the internet or in books, meaning that it has likely already been incorporated into a chatbot's training data. "There are probably thousands of places on webpages that describe a trip to Boston," Kathleen Creel, a professor of philosophy and computer science at Northeastern University, told me. There's people on Reddit talking about living in Boston and what they like."
Sora: OpenAI launches tool that instantly creates video from text
OpenAI revealed a tool on Thursday that can generate videos from text prompts. The new model, nicknamed Sora after the Japanese word for "sky", can produce realistic footage up to a minute long that adheres to a user's instructions on both subject matter and style. According to a company blogpost, the model is also able to create a video based on a still image or extend existing footage with new material. "We're teaching AI to understand and simulate the physical world in motion, with the goal of training models that help people solve problems that require real-world interaction," the blogpost reads. One video included among several initial examples from the company was based on the prompt: "A movie trailer featuring the adventures of the 30-year-old space man wearing a red wool knitted motorcycle helmet, blue sky, salt desert, cinematic style, shot on 35mm film, vivid colors."
OpenAI's new Sora model can generate minute-long videos from text prompts
OpenAI on Thursday announced Sora, a brand new model that generates high-definition videos up to one minute in length from text prompts. Sora, which means "sky" in Japanese, won't be available to the general public any time soon. Instead, OpenAI is making it available to a small group of academics and researchers who will assess harm and its potential for misuse. "Sora is able to generate complex scenes with multiple characters, specific types of motion, and accurate details of the subject and background," the company said on its website. "The model understands not only what the user has asked for in the prompt, but also how those things exist in the physical world."
OpenAI's Sora Turns AI Prompts Into Photorealistic Videos
We already know that OpenAI's chatbots can pass the bar exam without going to law school. Now, just in time for the Oscars, a new OpenAI app called Sora hopes to master cinema without going to film school. For now a research product, Sora is going out to a few select creators and a number of security experts who will red-team it for safety vulnerabilities. OpenAI plans to make it available to all wannabe auteurs at some unspecified date, but it decided to preview it in advance. Other companies, from giants like Google to startups like Runway, have already revealed text-to-video AI projects.
Google's new version of Gemini can handle far bigger amounts of data
The model was also able to identify moments of humor. When asked by the researchers to find a funny moment in the Apollo transcript, it picked out when astronaut Mike Collins referred to Armstrong as "the Czar." (Probably not the best line, but you get the point). In another demonstration, the team uploaded a 44-minute silent film featuring Buster Keaton and asked the AI to identify what information was on a piece of paper that, at some point in the movie, is removed from a character's pocket. In less than a minute, the model found the scene and correctly recalled the text written on the paper. Researchers also repeated a similar task from the Apollo experiment, asking the model to find a scene in the film based on a drawing, which it completed.
Improving Black-box Robustness with In-Context Rewriting
O'Brien, Kyle, Ng, Nathan, Puri, Isha, Mendez, Jorge, Palangi, Hamid, Kim, Yoon, Ghassemi, Marzyeh, Hartvigsen, Thomas
Machine learning models often excel on in-distribution (ID) data but struggle with unseen out-of-distribution (OOD) inputs. Most techniques for improving OOD robustness are not applicable to settings where the model is effectively a black box, such as when the weights are frozen, retraining is costly, or the model is leveraged via an API. Test-time augmentation (TTA) is a simple post-hoc technique for improving robustness that sidesteps black-box constraints by aggregating predictions across multiple augmentations of the test input. TTA has seen limited use in NLP due to the challenge of generating effective natural language augmentations. In this work, we propose LLM-TTA, which uses LLM-generated augmentations as TTA's augmentation function. LLM-TTA outperforms conventional augmentation functions across sentiment, toxicity, and news classification tasks for BERT and T5 models, with BERT's OOD robustness improving by an average of 4.30 percentage points without regressing average ID performance. We explore selectively augmenting inputs based on prediction entropy to reduce the rate of expensive LLM augmentations, allowing us to maintain performance gains while reducing the average number of generated augmentations by 57.76%. LLM-TTA is agnostic to the task model architecture, does not require OOD labels, and is effective across low and high-resource settings. We share our data, models, and code for reproducibility.