Large Language Model
Can AI Relate: Testing Large Language Model Response for Mental Health Support
Gabriel, Saadia, Puri, Isha, Xu, Xuhai, Malgaroli, Matteo, Ghassemi, Marzyeh
Large language models (LLMs) are already being piloted for clinical use in hospital systems like NYU Langone, Dana-Farber and the NHS. A proposed deployment use case is psychotherapy, where a LLM-powered chatbot can treat a patient undergoing a mental health crisis. Deployment of LLMs for mental health response could hypothetically broaden access to psychotherapy and provide new possibilities for personalizing care. However, recent high-profile failures, like damaging dieting advice offered by the Tessa chatbot to patients with eating disorders, have led to doubt about their reliability in high-stakes and safety-critical settings. In this work, we develop an evaluation framework for determining whether LLM response is a viable and ethical path forward for the automation of mental health treatment. Using human evaluation with trained clinicians and automatic quality-of-care metrics grounded in psychology research, we compare the responses provided by peer-to-peer responders to those provided by a state-of-the-art LLM. We show that LLMs like GPT-4 use implicit and explicit cues to infer patient demographics like race. We then show that there are statistically significant discrepancies between patient subgroups: Responses to Black posters consistently have lower empathy than for any other demographic group (2%-13% lower than the control group). Promisingly, we do find that the manner in which responses are generated significantly impacts the quality of the response. We conclude by proposing safety guidelines for the potential deployment of LLMs for mental health response.
AI chatbots' safeguards can be easily bypassed, say UK researchers
Guardrails to prevent artificial intelligence models behind chatbots from issuing illegal, toxic or explicit responses can be bypassed with simple techniques, UK government researchers have found. The UK's AI Safety Institute (AISI) said systems it had tested were "highly vulnerable" to jailbreaks, a term for text prompts designed to elicit a response that a model is supposedly trained to avoid issuing. The AISI said it had tested five unnamed large language models (LLM) – the technology that underpins chatbots – and circumvented their safeguards with relative ease, even without concerted attempts to beat their guardrails. "All tested LLMs remain highly vulnerable to basic jailbreaks, and some will provide harmful outputs even without dedicated attempts to circumvent their safeguards," wrote AISI researchers in an update on their testing regime. The AISI found that safeguards could be circumvented with "relatively simple" attacks, by, for instance, instructing the system to start its response with phrases like "Sure, I'm happy to help".
7 things Google just announced that are worth keeping a close eye on
ZeroEyes CEO Mike Lahiff joins'Fox & Friends' to explain how the technology works to help keep students safe in schools. Google's flagship developer conference called I/O just wrapped up with interesting leaps in how the Big Tech giant is planning to change the world. Here are the seven biggest things we learned from Google at I/O 2024. Google's I/O event was largely an opportunity for it to make its case to developers -- and, to a lesser extent, consumers -- as to why its artificial intelligence is ahead of rivals Microsoft and OpenAI. Here's a rundown of the seven highlights to keep an eye on.
Sam Altman is 'embarrassed' that OpenAI threatened to revoke equity if exiting employees wouldn't sign an NDA
OpenAI reportedly made exiting employees choose between keeping their vested equity and being able to speak out against the company. According to Vox, which viewed the document in question, employees could "lose all vested equity they earned during their time at the company, which is likely worth millions of dollars" if they didn't sign a nondisclosure and non-disparagement agreement, thanks to a provision in the off-boarding papers. OpenAI CEO Sam Altman confirmed in a tweet on Saturday evening that such a provision did exist, but said "we have never clawed back anyone's vested equity, nor will we do that if people do not sign a separation agreement (or don't agree to a non-disparagement agreement)." An OpenAI spokesperson echoed this in a statement to Vox, and Altman said the company "was already in the process of fixing the standard exit paperwork over the past month or so." But as Vox notes in its report, at least one former OpenAI employee has spoken publicly about sacrificing equity by declining to sign an NDA upon leaving.
Conversational Disease Diagnosis via External Planner-Controlled Large Language Models
Sun, Zhoujian, Luo, Cheng, Liu, Ziyi, Huang, Zhengxing
The development of large language models (LLMs) has brought unprecedented possibilities for artificial intelligence (AI) based medical diagnosis. However, the application perspective of LLMs in real diagnostic scenarios is still unclear because they are not adept at collecting patient data proactively. This study presents a LLM-based diagnostic system that enhances planning capabilities by emulating doctors. Our system involves two external planners to handle planning tasks. The first planner employs a reinforcement learning approach to formulate disease screening questions and conduct initial diagnoses. The second planner uses LLMs to parse medical guidelines and conduct differential diagnoses. By utilizing real patient electronic medical record data, we constructed simulated dialogues between virtual patients and doctors and evaluated the diagnostic abilities of our system. We demonstrated that our system obtained impressive performance in both disease screening and differential diagnoses tasks. This research represents a step towards more seamlessly integrating AI into clinical settings, potentially enhancing the accuracy and accessibility of medical diagnostics.
Generative Students: Using LLM-Simulated Student Profiles to Support Question Item Evaluation
Evaluating the quality of automatically generated question items has been a long standing challenge. In this paper, we leverage LLMs to simulate student profiles and generate responses to multiple-choice questions (MCQs). The generative students' responses to MCQs can further support question item evaluation. We propose Generative Students, a prompt architecture designed based on the KLI framework. A generative student profile is a function of the list of knowledge components the student has mastered, has confusion about or has no evidence of knowledge of. We instantiate the Generative Students concept on the subject domain of heuristic evaluation. We created 45 generative students using GPT-4 and had them respond to 20 MCQs. We found that the generative students produced logical and believable responses that were aligned with their profiles. We then compared the generative students' responses to real students' responses on the same set of MCQs and found a high correlation. Moreover, there was considerable overlap in the difficult questions identified by generative students and real students. A subsequent case study demonstrated that an instructor could improve question quality based on the signals provided by Generative Students.
Token-wise Influential Training Data Retrieval for Large Language Models
Lin, Huawei, Long, Jikai, Xu, Zhaozhuo, Zhao, Weijie
Given a Large Language Model (LLM) generation, how can we identify which training data led to this generation? In this paper, we proposed RapidIn, a scalable framework adapting to LLMs for estimating the influence of each training data. The proposed framework consists of two stages: caching and retrieval. First, we compress the gradient vectors by over 200,000x, allowing them to be cached on disk or in GPU/CPU memory. Then, given a generation, RapidIn efficiently traverses the cached gradients to estimate the influence within minutes, achieving over a 6,326x speedup. Moreover, RapidIn supports multi-GPU parallelization to substantially accelerate caching and retrieval. Our empirical result confirms the efficiency and effectiveness of RapidIn.
ColorFoil: Investigating Color Blindness in Large Vision and Language Models
Samin, Ahnaf Mozib, Ahmed, M. Firoz, Rafee, Md. Mushtaq Shahriyar
In this benchmark, With the utilization of Transformer architecture, large foils are generated from the existing V&L datasets for each Vision and Language (V&L) models have shown promising of the tasks. A foil is referred to as a distractor or slightly performance in even zero-shot settings. Several studies, incorrect example that is passed along with the correct example however, indicate a lack of robustness of the models when to the V&L model to assess the model's ability to dealing with complex linguistics and visual attributes. In correctly distinguish them [17, 22]. Although the existing this work, we introduce a novel V&L benchmark - Color-V&L benchmarks like VALSE help the community to test Foil, by creating color-related foils to assess the models' the capabilities of V&L models, there is still much work to perception ability to detect colors like red, white, green, etc. be done to evaluate the robustness and generalizability of We evaluate seven state-of-the-art V&L models including the models on numerous other tasks. It remains unknown CLIP, ViLT, GroupViT, and BridgeTower, etc. in a zero-shot how well the large V&L models can perceive colors from setting and present intriguing findings from the V&L models.
Exploring the Capabilities of Prompted Large Language Models in Educational and Assessment Applications
Maity, Subhankar, Deroy, Aniket, Sarkar, Sudeshna
In the era of generative artificial intelligence (AI), the fusion of large language models (LLMs) offers unprecedented opportunities for innovation in the field of modern education. We embark on an exploration of prompted LLMs within the context of educational and assessment applications to uncover their potential. Through a series of carefully crafted research questions, we investigate the effectiveness of prompt-based techniques in generating open-ended questions from school-level textbooks, assess their efficiency in generating open-ended questions from undergraduate-level technical textbooks, and explore the feasibility of employing a chain-of-thought inspired multi-stage prompting approach for language-agnostic multiple-choice question (MCQ) generation. Additionally, we evaluate the ability of prompted LLMs for language learning, exemplified through a case study in the low-resource Indian language Bengali, to explain Bengali grammatical errors. We also evaluate the potential of prompted LLMs to assess human resource (HR) spoken interview transcripts. By juxtaposing the capabilities of LLMs with those of human experts across various educational tasks and domains, our aim is to shed light on the potential and limitations of LLMs in reshaping educational practices.
Zero-Shot Stance Detection using Contextual Data Generation with LLMs
Mahmoudi, Ghazaleh, Behkamkia, Babak, Eetemadi, Sauleh
Stance detection, the classification of attitudes expressed in a text towards a specific topic, is vital for applications like fake news detection and opinion mining. However, the scarcity of labeled data remains a challenge for this task. To address this problem, we propose Dynamic Model Adaptation with Contextual Data Generation (DyMoAdapt) that combines Few-Shot Learning and Large Language Models. In this approach, we aim to fine-tune an existing model at test time. We achieve this by generating new topic-specific data using GPT-3. This method could enhance performance by allowing the adaptation of the model to new topics. However, the results did not increase as we expected. Furthermore, we introduce the Multi Generated Topic VAST (MGT-VAST) dataset, which extends VAST using GPT-3. In this dataset, each context is associated with multiple topics, allowing the model to understand the relationship between contexts and various potential topics