Toward the Evaluation of Large Language Models Considering Score Variance across Instruction Templates

Open in new window