Can Large Language Models Reason? A Characterization via 3-SAT

Hazra, Rishi, Venturato, Gabriele, Martires, Pedro Zuidberg Dos, De Raedt, Luc

arXiv.org Artificial Intelligence 

Large Language Models (LLMs) have been touted as AI models possessing advanced reasoning abilities. However, recent works have shown that LLMs often bypass true reasoning using shortcuts, sparking skepticism. To study the reasoning capabilities in a principled fashion, we adopt a computational theory perspective and propose an experimental protocol centered on 3-SAT - the prototypical NPcomplete problem lying at the core of logical reasoning and constraint satisfaction tasks. Specifically, we examine the phase transitions in random 3-SAT and characterize the reasoning abilities of LLMs by varying the inherent hardness of the problem instances. Our experimental evidence shows that LLMs are incapable of performing true reasoning, as required for solving 3-SAT problems. Moreover, we observe significant performance variation based on the inherent hardness of the problems - performing poorly on harder instances and vice versa. Importantly, we show that integrating external reasoners can considerably enhance LLM performance. By following a principled experimental protocol, our study draws concrete conclusions and moves beyond the anecdotal evidence often found in LLM reasoning research. The success and versatility of Large Language Models (LLMs) have sparked widespread interest and debate on whether LLMs are capable of reasoning. The answer to this question may depend on the perspective on reasoning one takes, whether it is more oriented toward common sense reasoning (Davis & Marcus, 2015) or towards logical or deductive reasoning (Genesereth & Nilsson, 1987). We will adhere to Leon Bottou's definition, which defines reasoning as "algebraically manipulating previously acquired knowledge in order to answer a new question" (Bottou, 2014). This is aligned with Russell and Norvig's description of artificial intelligence as rational thinking (Russell & Norvig, 2010). Recent studies suggest that LLMs are inherently capable of zero-shot reasoning (Kojima et al., 2022) (i.e. This ability has been shown to emerge and improve with scale (Wei et al., 2022a; Srivastava et al., 2023), and can be further enhanced by using smart prompting techniques that encourage LLMs to think stepby-step (Kojima et al., 2022; Wei et al., 2022b; Zhou et al., 2023; Yao et al., 2023b; Prasad et al., 2023).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found