Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi

Kaptur, Dandan Chen, Huang, Yue, Ji, Xuejun Ryan, Guo, Yanhui, Kaptur, Bradley

arXiv.org Artificial Intelligence 

We evaluated their performance by comparing LLM-generated codes with human-generated codes from a peer-reviewed systematic review on assessment. Our findings suggested that LLMs' performance fluctuates by data volume and question complexity for systematic reviews. Word count: 785 Introduction Despite the growing use of Large Language Models (LLMs) in academic research, questions about their accuracy and reliability persist. This study compares GPT-4, a widely used LLM, and Kimi, less studied particularly in Western academia. We conducted a comparative analysis of GPT-4 and Kimi, to explore nuances in their performance for different situations in the context of text coding for systematic reviews.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found