Large Language Model
Citation-Enhanced Generation for LLM-based Chatbots
Li, Weitao, Li, Junkai, Ma, Weizhi, Liu, Yang
Large language models (LLMs) exhibit powerful general intelligence across diverse scenarios, including their integration into chatbots. However, a vital challenge of LLM-based chatbots is that they may produce hallucinated content in responses, which significantly limits their applicability. Various efforts have been made to alleviate hallucination, such as retrieval augmented generation and reinforcement learning with human feedback, but most of them require additional training and data annotation. In this paper, we propose a novel post-hoc Citation-Enhanced Generation (CEG) approach combined with retrieval argumentation. Unlike previous studies that focus on preventing hallucinations during generation, our method addresses this issue in a post-hoc way. It incorporates a retrieval module to search for supporting documents relevant to the generated content, and employs a natural language inference-based citation generation module. Once the statements in the generated content lack of reference, our model can regenerate responses until all statements are supported by citations. Note that our method is a training-free plug-and-play plugin that is capable of various LLMs. Experiments on various hallucination-related datasets show our framework outperforms state-of-the-art methods in both hallucination detection and response regeneration on three benchmarks. Our codes and dataset will be publicly available.
AI's craving for data is matched only by a runaway thirst for water and energy John Naughton
One of the most pernicious myths about digital technology is that it is somehow weightless or immaterial. Remember all that early talk about the "paperless" office and "frictionless" transactions? And of course, while our personal electronic devices do use some electricity, compared with the washing machine or the dishwasher, it's trivial. Belief in this comforting story, however, might not survive an encounter with Kate Crawford's seminal book, Atlas of AI, or the striking Anatomy of an AI System graphic she composed with Vladan Joler. And it certainly wouldn't survive a visit to a datacentre โ one of those enormous metallic sheds housing tens or even hundreds of thousands of servers humming away, consuming massive amounts of electricity and needing lots of water for their cooling systems.
These Companies Have a Plan to Kill Apps
Everyone wants to kill the app. There's a wave of companies building so-called app-less phones and gadgets, leveraging artificial intelligence advancements to create smarter virtual assistants that can handle all kinds of tasks through one portal, bypassing the need for specific apps for a particular function. We might be witnessing the early stages of the first major smartphone evolution since the introduction of the iPhone--or an AI-hype-fueled gimmick. There's the Humane Ai Pin, a wearable that can identify objects, take photos, and project information into the palm of your hand. It's powered by a digital assistant that uses multiple large language models, such as ChatGPT, and it's designed to reduce reliance on the smartphone.
Elon Musk's OpenAI Lawsuit: Corporate Conniving or Battle for Humankind?
This week, Felix Salmon, Emily Peck and Elizabeth Spiers ponder the future of computers, cars, andโฆfast food? They discuss why Elon Musk is suing Sam Altman and OpenAI and the altruistic origins of ChatGPT. Also: Wendy's "surge pricing" gaff had customers crying foul and Apple's electric car has been scrapped. If you enjoy this show, please consider signing up for Slate Plus. Slate Plus members get an ad-free experience across the network and an additional segment of our show every week.
Elon Musk sues OpenAI and Sam Altman for violating the company's principles
OpenAI, the influential artificial intelligence company that ousted and then reinstated its high-profile CEO three months ago, faces a new drama: a lawsuit from Elon Musk, one of the richest men in the world and a co-founder of the AI lab. Musk sued OpenAI and its CEO, Sam Altman, accusing them of breaching a contract by putting profits and commercial interests in developing AI ahead of the public good. A multibillion-dollar partnership that OpenAI developed with Microsoft, Musk said, represented an abandonment of a founding pledge to carefully develop AI and make the technology publicly available. "OpenAI has been transformed into a closed-source de facto subsidiary of the largest technology company, Microsoft," said the lawsuit filed Thursday in Superior Court in San Francisco.
Nvidia CEO says AI could pass human tests in five years
Nvidia Chief Executive Jensen Huang on Friday said that artificial general intelligence could -- by some definitions -- arrive in as little as five years. Huang, who heads the world's leading maker of artificial intelligence chips used to create systems like OpenAI's ChatGPT, was responding to a question at an economic forum held at Stanford University about how long it would take to achieve one of Silicon Valley's long-held goals of creating computers that can think like humans. Huang said that the answer largely depends on how the goal is defined. If the definition is the ability to pass human tests, Huang said, artificial general intelligence (AGI) will arrive soon.
Improving the Validity of Automatically Generated Feedback via Reinforcement Learning
Scarlatos, Alexander, Smith, Digory, Woodhead, Simon, Lan, Andrew
Automatically generating feedback via large language models (LLMs) in intelligent tutoring systems and online learning platforms has the potential to improve the learning outcomes of many students. However, both feedback generation and evaluation are challenging: feedback content has to be valid especially in subjects like math, which requires models to understand the problem, the solution, and where the student's error lies. Feedback also has to be pedagogically valid to reflect effective tutoring strategies, such as explaining possible misconceptions and encouraging the student, among other desirable features. In this work, we address both problems of automatically generating and evaluating feedback while considering both correctness and alignment. First, we propose a rubric for evaluating math feedback and show that GPT-4 is able to effectively use it to annotate human-written and LLM-generated feedback. Second, we propose a framework for feedback generation that optimizes both correctness and alignment using reinforcement learning (RL). Specifically, we use GPT-4's annotations to create preferences over feedback pairs in an augmented dataset for training via direct preference optimization (DPO). We show that our methods significantly increase the correctness and alignment of generated feedback with Llama 2, an open-source LLM, qualitatively analyze our generation and evaluation systems using case studies, and outline several areas for future work.
SyllabusQA: A Course Logistics Question Answering Dataset
Fernandez, Nigel, Scarlatos, Alexander, Lan, Andrew
Moreover, text similarity metrics may not be suitable In educational applications, artificial intelligence for some open-ended natural language generation (AI) approaches have shown significant promise in tasks (Amidei et al., 2018). As an example, the answer improving learning outcomes (Aleven et al., 2016; "The final exam will be on Dec 15", has high VanLehn, 2011), by automatically providing feedback surface-level textual similarity with the reference to students or engaging in tutoring dialogues answer, "The final exam is on Dec 14", but contains with them. The key idea is to use AI to create an ondemand a critical factual error that may lead to significant virtual teaching assistant to interact with negative consequences to students. Meanwhile, many students simultaneously; see, e.g., Khamigo human instructors and teaching assistants often answer from Khan Academy (Academy, 2022). These approaches student questions in a concise way, without can scale up the effort of expert human giving any unnecessary information. Therefore, it teachers and tutors, and relieve them from doing is important for LLM-based approaches to generate repetitive tasks so that they can focus on providing answers that are both concise and precise.
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
Benster, Tyler, Wilson, Guy, Elisha, Reshef, Willett, Francis R, Druckmann, Shaul
Silent Speech Interfaces (SSIs) offer a noninvasive alternative to brain-computer interfaces for soundless verbal communication. We introduce Multimodal Orofacial Neural Audio (MONA), a system that leverages cross-modal alignment through novel loss functions--cross-contrast (crossCon) and supervised temporal contrast (supTcon)--to train a multimodal model with a shared latent representation. This architecture enables the use of audio-only datasets like LibriSpeech to improve silent speech recognition. Additionally, our introduction of Large Language Model (LLM) Integrated Scoring Adjustment (LISA) significantly improves recognition accuracy. Together, MONA LISA reduces the state-of-the-art word error rate (WER) from 28.8% to 12.2% in the Gaddy (2020) benchmark dataset for silent speech on an open vocabulary. For vocal EMG recordings, our method improves the state-of-the-art from 23.3% to 3.7% WER. In the Brain-to-Text 2024 competition, LISA performs best, improving the top WER from 9.8% to 8.9%. To the best of our knowledge, this work represents the first instance where noninvasive silent speech recognition on an open vocabulary has cleared the threshold of 15% WER, demonstrating that SSIs can be a viable alternative to automatic speech recognition (ASR). Our work not only narrows the performance gap between silent and vocalized speech but also opens new possibilities in human-computer interaction, demonstrating the potential of cross-modal approaches in noisy and data-limited regimes.
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
Zeng, Yifan, Wu, Yiran, Zhang, Xiao, Wang, Huazheng, Wu, Qingyun
Despite extensive pre-training and fine-tuning in moral alignment to prevent generating harmful information at user request, large language models (LLMs) remain vulnerable to jailbreak attacks. In this paper, we propose AutoDefense, a response-filtering based multi-agent defense framework that filters harmful responses from LLMs. This framework assigns different roles to LLM agents and employs them to complete the defense task collaboratively. The division in tasks enhances the overall instruction-following of LLMs and enables the integration of other defense components as tools. AutoDefense can adapt to various sizes and kinds of open-source LLMs that serve as agents. Through conducting extensive experiments on a large scale of harmful and safe prompts, we validate the effectiveness of the proposed AutoDefense in improving the robustness against jailbreak attacks, while maintaining the performance at normal user request. Our code and data are publicly available at https://github.com/XHMY/AutoDefense.