accuracy and consistency
Generative AI for Research Data Processing: Lessons Learnt From Three Use Cases
Mitra, Modhurita, de Vos, Martine G., Cortinovis, Nicola, Ometto, Dawa
--There has been enormous interest in generative AI since ChatGPT was launched in 2022. However, there are concerns about the accuracy and consistency of the outputs of generative AI. We have carried out an exploratory study on the application of this new technology in research data processing. We identified tasks for which rule-based or traditional machine learning approaches were difficult to apply, and then performed these tasks using generative AI. We demonstrate the feasibility of using the generative AI model Claude 3 Opus in three research projects involving complex data processing tasks: 1) Information extraction: We extract plant species names from historical seedlists (catalogues of seeds) published by botanical gardens. We share the lessons we learnt from these use cases: How to determine if generative AI is an appropriate tool for a given data processing task, and if so, how to maximise the accuracy and consistency of the results obtained. In this paper, we share our insights on the application of generative AI in research software engineering projects. Generative AI can potentially be used to perform a wide variety of research data processing tasks, such as interpreting documents, extracting information from them, and classifying text into categories. Since the tasks are specified through prompts in natural language, the barrier to entry is low. Therefore, this tool can be easily used by domain experts in a wide range of fields, with varying levels of programming skills and depth of knowledge of technical topics such as machine learning.
Logical Consistency of Large Language Models in Fact-checking
Ghosh, Bishwamittra, Hasan, Sarah, Arafat, Naheed Anjum, Khan, Arijit
In recent years, large language models (LLMs) have demonstrated significant success in performing varied natural language tasks such as language translation, question-answering, summarizing, fact-checking, etc. Despite LLMs' impressive ability to generate human-like texts, LLMs are infamous for their inconsistent responses -- a meaning-preserving change in the input query results in an inconsistent response and attributes to vulnerabilities of LLMs such as hallucination, jailbreaking, etc. Consequently, existing research focuses on simple paraphrasing-based consistency assessment of LLMs, and ignores complex queries that necessitates an even better understanding of logical reasoning by an LLM. Our work therefore addresses the logical inconsistency of LLMs under complex logical queries with primitive logical operators, e.g., negation, conjunction, and disjunction. As a test bed, we consider retrieval-augmented LLMs on a fact-checking task involving propositional logic queries from real-world knowledge graphs (KGs). Our contributions are three-fold. Benchmark: We introduce three logical fact-checking datasets over KGs for community development towards logically consistent LLMs. Assessment: We propose consistency measures of LLMs on propositional logic queries as input and demonstrate that existing LLMs lack logical consistency, specially on complex queries. Improvement: We employ supervised fine-tuning to improve the logical consistency of LLMs on the complex fact-checking task with KG contexts.
Accuracy and Consistency of LLMs in the Registered Dietitian Exam: The Impact of Prompt Engineering and Knowledge Retrieval
Azimi, Iman, Qi, Mohan, Wang, Li, Rahmani, Amir M., Li, Youlin
Large language models (LLMs) are fundamentally transforming human-facing applications in the health and well-being domains: boosting patient engagement, accelerating clinical decision-making, and facilitating medical education. Although state-of-the-art LLMs have shown superior performance in several conversational applications, evaluations within nutrition and diet applications are still insufficient. In this paper, we propose to employ the Registered Dietitian (RD) exam to conduct a standard and comprehensive evaluation of state-of-the-art LLMs, GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro, assessing both accuracy and consistency in nutrition queries. Our evaluation includes 1050 RD exam questions encompassing several nutrition topics and proficiency levels. In addition, for the first time, we examine the impact of Zero-Shot (ZS), Chain of Thought (CoT), Chain of Thought with Self Consistency (CoT-SC), and Retrieval Augmented Prompting (RAP) on both accuracy and consistency of the responses. Our findings revealed that while these LLMs obtained acceptable overall performance, their results varied considerably with different prompts and question domains. GPT-4o with CoT-SC prompting outperformed the other approaches, whereas Gemini 1.5 Pro with ZS recorded the highest consistency. For GPT-4o and Claude 3.5, CoT improved the accuracy, and CoT-SC improved both accuracy and consistency. RAP was particularly effective for GPT-4o to answer Expert level questions. Consequently, choosing the appropriate LLM and prompting technique, tailored to the proficiency level and specific domain, can mitigate errors and potential risks in diet and nutrition chatbots.
The Application of ChatGPT in Responding to Questions Related to the Boston Bowel Preparation Scale
Liu, Xiaoqiang, Wang, Yubin, Huang, Zicheng, Xu, Boming, Zeng, Yilin, Chen, Xinqi, Wang, Zilong, Yang, Enning, Lei, Xiaoxuan, Huang, Yisen, Liu, Xiaobo
Background: Colonoscopy, a crucial diagnostic tool in gastroenterology, depends heavily on superior bowel preparation. ChatGPT, a large language model with emergent intelligence which also exhibits potential in medical applications. This study aims to assess the accuracy and consistency of ChatGPT in using the Boston Bowel Preparation Scale (BBPS) for colonoscopy assessment. Methods: We retrospectively collected 233 colonoscopy images from 2020 to 2023. These images were evaluated using the BBPS by 3 senior endoscopists and 3 novice endoscopists. Additionally, ChatGPT also assessed these images, having been divided into three groups and undergone specific Fine-tuning. Consistency was evaluated through two rounds of testing. Results: In the initial round, ChatGPT's accuracy varied between 48.93% and 62.66%, trailing the endoscopists' accuracy of 76.68% to 77.83%. Kappa values for ChatGPT was between 0.52 and 0.53, compared to 0.75 to 0.87 for the endoscopists. Conclusion: While ChatGPT shows promise in bowel preparation scoring, it currently does not match the accuracy and consistency of experienced endoscopists. Future research should focus on in-depth Fine-tuning.
Semantic Consistency for Assuring Reliability of Large Language Models
Raj, Harsh, Gupta, Vipul, Rosati, Domenic, Majumdar, Subhabrata
Large Language Models (LLMs) exhibit remarkable fluency and competence across various natural language tasks. However, recent research has highlighted their sensitivity to variations in input prompts. To deploy LLMs in a safe and reliable manner, it is crucial for their outputs to be consistent when prompted with expressions that carry the same meaning or intent. While some existing work has explored how state-of-the-art LLMs address this issue, their evaluations have been confined to assessing lexical equality of single- or multi-word answers, overlooking the consistency of generative text sequences. For a more comprehensive understanding of the consistency of LLMs in open-ended text generation scenarios, we introduce a general measure of semantic consistency, and formulate multiple versions of this metric to evaluate the performance of various LLMs. Our proposal demonstrates significantly higher consistency and stronger correlation with human evaluations of output consistency than traditional metrics based on lexical consistency. Finally, we propose a novel prompting strategy, called Ask-to-Choose (A2C), to enhance semantic consistency. When evaluated for closed-book question answering based on answer variations from the TruthfulQA benchmark, A2C increases accuracy metrics for pretrained and finetuned LLMs by up to 47%, and semantic consistency metrics for instruction-tuned models by up to 7-fold.
Artificial intelligence and automation: Examples, benefits and more
As artificial intelligence and automation technologies continue to improve, they will become more important in driving the growth of new industries that are based on data. Artificial intelligence is the development of computer systems that are able to perform tasks that typically require human intelligence, such as visual perception, speech recognition, decision-making, and problem-solving. AI systems are typically designed to be able to learn from experience, adapt to new inputs, and improve their performance over time. Automation, on the other hand, refers to the use of technology to automate tasks that were previously performed by humans. This can include everything from simple tasks like data entry to more complex tasks like driving a car or managing a supply chain. Automation can be powered by a variety of technologies, including AI, robotics, and machine learning.
Improves Document AI accuracy and consistency with EKG
Imagine a bank employee entering this to capture a customer's income to qualify them for a loan. She might take the time to go into Google and find the right correct legal entity name; or she might just guess and move on, potentially creating data reconciliation headaches down the line. With Document AI, there is a better way. Our native integration with the knowledge graph means that we can deliver both the specific text from the payslip as well as Google's best understanding of the actual name of the company that operates at this address. This is an important step to translate from "what has been said" on a document to "what does it mean".
How to Ensure Data Quality for AI - insideBIGDATA
In this special guest feature, Wilson Pang, CTO of Appen, offers a few quality controls that organizations can implement to allow for the most accurate and consistent data annotation process possible. Wilson joined Appen in November 2018 and is responsible for the company's products and technology. Wilson has over seventeen years' experience in software engineering and data science. Prior to joining Appen, Wilson was Chief Data Officer of CTrip in China, the second largest online travel agency company in the world where he led data engineers, analysts, data product managers and scientists to improve user experience and increase operational efficiency that grew the business. Before that, he was senior director of engineering in eBay in California and provided leadership to various domains including data service and solutions, search science, marketing technology and billing systems. Wilson obtained his Masters and Bachelor's degrees of Electric Engineering from Zhejiang University in China.
How to Ensure Data Quality for AI - insideBIGDATA
In this special guest feature, Wilson Pang, CTO of Appen, offers a few quality controls that organizations can implement to allow for the most accurate and consistent data annotation process possible. Wilson joined Appen in November 2018 and is responsible for the company's products and technology. Wilson has over seventeen years' experience in software engineering and data science. Prior to joining Appen, Wilson was Chief Data Officer of CTrip in China, the second largest online travel agency company in the world where he led data engineers, analysts, data product managers and scientists to improve user experience and increase operational efficiency that grew the business. Before that, he was senior director of engineering in eBay in California and provided leadership to various domains including data service and solutions, search science, marketing technology and billing systems. Wilson obtained his Masters and Bachelor's degrees of Electric Engineering from Zhejiang University in China.
Successful AI Development Means Fielding The Best Team
Breakthroughs in AI happen with a team of people working behind the scenes to make it all possible. If AI development were a sport, it'd be closer to baseball than boxing. Headlines might make it seem like AI breakthroughs happen with a big knockout punch, but the reality is more akin to a baseball team grinding through a 162-game season. It's a process that involves having the right people in place over a long stretch, and fielding the best team is essential for success. The root cause for this can be found in the data.