Assessing Programming Task Difficulty for Efficient Evaluation of Large Language Models

Open in new window