Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi
Kaptur, Dandan Chen, Huang, Yue, Ji, Xuejun Ryan, Guo, Yanhui, Kaptur, Bradley
–arXiv.org Artificial Intelligence
We evaluated their performance by comparing LLM-generated codes with human-generated codes from a peer-reviewed systematic review on assessment. Our findings suggested that LLMs' performance fluctuates by data volume and question complexity for systematic reviews. Word count: 785 Introduction Despite the growing use of Large Language Models (LLMs) in academic research, questions about their accuracy and reliability persist. This study compares GPT-4, a widely used LLM, and Kimi, less studied particularly in Western academia. We conducted a comparative analysis of GPT-4 and Kimi, to explore nuances in their performance for different situations in the context of text coding for systematic reviews.
arXiv.org Artificial Intelligence
Apr-30-2025