Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
Gong, Linyuan, Wang, Sida, Elhoushi, Mostafa, Cheung, Alvin
–arXiv.org Artificial Intelligence
SAFIM emphasizes Language Models (LLMs) on the code Fill-inthe-Middle syntax-aware completion within code's Abstract Syntax (FIM) task. This benchmark focuses Tree (AST), targeting algorithmic blocks, control-flow expressions, on syntax-aware completions of program structures and API function calls, unlike existing Fillin-the such as code blocks and conditional expressions, Middle (FIM) benchmarks such as HumanEval-and includes 17,720 examples from multiple Infilling (Bavarian et al., 2022), which are based on filling programming languages, sourced from recent randomly masked lines or character spans. SAFIM is code submissions after April 2022 to minimize sourced from code on Codeforces and GitHub created after data contamination. SAFIM provides a April 2022, deliberately aiming to avoid overlap with mainstream robust framework with various prompt designs open-source pretraining corpora like The Stack (Kocetkov and novel syntax-aware post-processing techniques, et al., 2022). This approach reduces the risks of facilitating accurate and fair comparisons data contamination caused by memoization of test cases, across LLMs.
arXiv.org Artificial Intelligence
Jun-22-2024
- Country:
- North America > United States
- California (0.04)
- Europe > Austria
- Vienna (0.14)
- North America > United States
- Genre:
- Research Report > New Finding (0.93)
- Technology: