Benchmarking Large Language Model Uncertainty for Prompt Optimization

Guo, Pei-Fu, Tsai, Yun-Da, Lin, Shou-De

arXiv.org Artificial Intelligence 

Prompt optimization algorithms for Large Language Models (LLMs) excel in multi-step reasoning but still lack effective uncertainty estimation. This paper introduces a benchmark dataset to evaluate uncertainty metrics, focusing on Answer, Correctness, Aleatoric, and Epistemic Uncertainty.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found