Goto

Collaborating Authors

 Large Language Model


VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models

arXiv.org Artificial Intelligence

Recent approaches in domain-specific named entity recognition (NER), such as biomedical NER, have shown remarkable advances. However, they still lack of faithfulness, producing erroneous predictions. We assume that knowledge of entities can be useful in verifying the correctness of the predictions. Despite the usefulness of knowledge, resolving such errors with knowledge is nontrivial, since the knowledge itself does not directly indicate the ground-truth label. To this end, we propose VerifiNER, a post-hoc verification framework that identifies errors from existing NER methods using knowledge and revises them into more faithful predictions. Our framework leverages the reasoning abilities of large language models to adequately ground on knowledge and the contextual information in the verification process. We validate effectiveness of VerifiNER through extensive experiments on biomedical datasets. The results suggest that VerifiNER can successfully verify errors from existing models as a model-agnostic approach. Further analyses on out-of-domain and low-resource settings show the usefulness of VerifiNER on real-world applications.


RewardBench: Evaluating Reward Models for Language Modeling

arXiv.org Artificial Intelligence

Reward models (RMs) are at the crux of successfully using RLHF to align pretrained models to human preferences, yet there has been relatively little study that focuses on evaluation of those models. Evaluating reward models presents an opportunity to understand the opaque technologies used for alignment of language models and which values are embedded in them. Resources for reward model training and understanding are sparse in the nascent open-source community around them. To enhance scientific understanding of reward models, we present RewardBench, a benchmark dataset and code-base for evaluation. The RewardBench dataset is a collection of prompt-chosen-rejected trios spanning chat, reasoning, and safety, to benchmark how reward models perform on challenging, structured and out-of-distribution queries. We create specific comparison datasets for RMs that have subtle, but verifiable reasons (e.g. bugs, incorrect facts) why one answer should be preferred to another. On the RewardBench leaderboard, we evaluate reward models trained with a variety of methods, such as the direct MLE training of classifiers and the implicit reward modeling of Direct Preference Optimization (DPO). We present many findings on propensity for refusals, reasoning limitations, and instruction following shortcomings of various reward models towards a better understanding of the RLHF process.


Guiding Clinical Reasoning with Large Language Models via Knowledge Seeds

arXiv.org Artificial Intelligence

Clinical reasoning refers to the cognitive process that physicians employ in evaluating and managing patients. This process typically involves suggesting necessary examinations, diagnosing patients' diseases, and deciding on appropriate therapies, etc. Accurate clinical reasoning requires extensive medical knowledge and rich clinical experience, setting a high bar for physicians. This is particularly challenging in developing countries due to the overwhelming number of patients and limited physician resources, contributing significantly to global health inequity and necessitating automated clinical reasoning approaches. Recently, the emergence of large language models (LLMs) such as ChatGPT and GPT-4 have demonstrated their potential in clinical reasoning. However, these LLMs are prone to hallucination problems, and the reasoning process of LLMs may not align with the clinical decision path of physicians. In this study, we introduce a novel framework, In-Context Padding (ICP), designed to enhance LLMs with medical knowledge. Specifically, we infer critical clinical reasoning elements (referred to as knowledge seeds) and use these as anchors to guide the generation process of LLMs. Experiments on two clinical question datasets demonstrate that ICP significantly improves the clinical reasoning ability of LLMs.


Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models

arXiv.org Artificial Intelligence

Accurate and interpretable user satisfaction estimation (USE) is critical for understanding, evaluating, and continuously improving conversational systems. Users express their satisfaction or dissatisfaction with diverse conversational patterns in both general-purpose (ChatGPT and Bing Copilot) and task-oriented (customer service chatbot) conversational systems. Existing approaches based on featurized ML models or text embeddings fall short in extracting generalizable patterns and are hard to interpret. In this work, we show that LLMs can extract interpretable signals of user satisfaction from their natural language utterances more effectively than embedding-based approaches. Moreover, an LLM can be tailored for USE via an iterative prompting framework using supervision from labeled examples. The resulting method, Supervised Prompting for User satisfaction Rubrics (SPUR), not only has higher accuracy but is more interpretable as it scores user satisfaction via learned rubrics with a detailed breakdown.


Pearl: A Review-driven Persona-Knowledge Grounded Conversational Recommendation Dataset

arXiv.org Artificial Intelligence

Conversational recommender system is an emerging area that has garnered an increasing interest in the community, especially with the advancements in large language models (LLMs) that enable diverse reasoning over conversational input. Despite the progress, the field has many aspects left to explore. The currently available public datasets for conversational recommendation lack specific user preferences and explanations for recommendations, hindering high-quality recommendations. To address such challenges, we present a novel conversational recommendation dataset named PEARL, synthesized with persona- and knowledge-augmented LLM simulators. We obtain detailed persona and knowledge from real-world reviews and construct a large-scale dataset with over 57k dialogues. Our experimental results demonstrate that utterances in PEARL include more specific user preferences, show expertise in the target domain, and provide recommendations more relevant to the dialogue context than those in prior datasets.


k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text

arXiv.org Artificial Intelligence

Recent watermarked generation algorithms inject detectable signatures during language generation to facilitate post-hoc detection. While token-level watermarks are vulnerable to paraphrase attacks, SemStamp (Hou et al., 2023) applies watermark on the semantic representation of sentences and demonstrates promising robustness. SemStamp employs locality-sensitive hashing (LSH) to partition the semantic space with arbitrary hyperplanes, which results in a suboptimal tradeoff between robustness and speed. We propose k-SemStamp, a simple yet effective enhancement of SemStamp, utilizing k-means clustering as an alternative of LSH to partition the embedding space with awareness of inherent semantic structure. Experimental results indicate that k-SemStamp saliently improves its robustness and sampling efficiency while preserving the generation quality, advancing a more effective tool for machine-generated text detection.


A Timeline of All the Recent Accusations Leveled at OpenAI and Sam Altman

TIME - Tech

Recent weeks have not been kind to OpenAI. The release of the company's latest model, GPT-4o, has been somewhat overshadowed by a series of accusations leveled at both the company and its CEO, Sam Altman. This comes at the same time that several high-profile employees, including co-founder and chief scientist Ilya Sutskever, have chosen to leave the company. This is not the first time the Silicon Valley startup has been embroiled in scandal. In November, Altman was briefly ousted from the company after the board found he had not been "consistently candid" with them.


This Is What It Looks Like When AI Eats the World

The Atlantic - Technology

Tech evangelists like to say that AI will eat the world--a reference to a famous line about software from the venture capitalist Marc Andreessen. In the past few weeks, we've finally gotten a sense of what they mean. This spring, tech companies have made clear that AI will be a defining feature of online life, whether people want it to be or not. First, Meta surprised users with an AI chatbot that lives in the search bar on Instagram and Facebook. It has since informed European users that their data are being used to train its AI--presumably sent only to comply with the continent's privacy laws. OpenAI released GPT-4o, billed as a new, more powerful and conversational version of its large language model.


Don't Let Mistrust of Tech Companies Blind You to the Power of AI

WIRED

It seems evident to me that almost 70 years after the first conference on artificial intelligence--where the nascent field's leaders suggested the task would be completed within a decade--the field is now poised to make a transformational impact on our lives. We don't need to reach artificial general intelligence, or AGI, whatever that means, for this to happen. I wrote as much in this column three weeks ago, citing evidence that after the astonishing leap of large language models that gave us ChatGPT, the advancements had not "plateaued" as some critics were charging. I also disagreed with the wave of skeptics claiming that what looked amazing in OpenAI's GPT-4, Anthropic's Claude 3, Meta's Llama 3, and an armada of Microsoft Copilots was merely a linguistic variation of a card trick. The hype, I insisted, is justified.


Engadget Podcast: How AI will shape Apple's WWDC 2024

Engadget

We're gearing up to cover Apple's Worldwide Developers Conference (WWDC) next week! In this episode, Cherlynn and Devindra dive into everything they expect at WWDC: Tons of AI announcements; more on iOS 18, iPadOS 18, and macOS 15; and hopefully some improvements for Vision Pro and visionOS. In addition, we chat about what we expect to see at Summer Game Fest and demonstrate how we used an AI editing tool to clear up some awful podcast audio. Devindra also talks with Justin Samuels, the founder of Render ATL, about why he started a massive tech conference in Atlanta. Listen below or subscribe on your podcast app of choice. If you've got suggestions or topics you'd like covered on the show, be sure to email us or drop a note in the comments! And be sure to check out our other podcast, Engadget News! Humane AI warns users its battery case "may pose a fire risk" – 34:36 Welcome back to the Engadget podcast. This week we are getting ready for WWDC 2024 happening in a couple of days.