Assessing Code Understanding in LLMs
Laneve, Cosimo, Spanò, Alvise, Ressi, Dalila, Rossi, Sabina, Bugliesi, Michele
–arXiv.org Artificial Intelligence
We present an empirical evaluation of Large Language Models in code understanding associated with non-trivial, semantic-preserving program transformations such as copy propagation or constant folding. Our findings show that LLMs fail to judge semantic equivalence in approximately 41\% of cases when no context is provided and in 29\% when given a simple generic context. To improve accuracy, we advocate integrating LLMs with code-optimization tools to enhance training and facilitate more robust program understanding.
arXiv.org Artificial Intelligence
Mar-31-2025
- Country:
- North America > United States
- Massachusetts > Suffolk County > Boston (0.04)
- Europe > Italy
- Veneto > Venice (0.04)
- Emilia-Romagna > Metropolitan City of Bologna
- Bologna (0.04)
- North America > United States
- Genre:
- Research Report > New Finding (0.86)
- Industry:
- Information Technology > Security & Privacy (0.46)
- Technology: