Aiming to address this limitation, we present Easy2Hard-Bench, a consistently formatted collection of 6 benchmark datasets spanning various domains, such as mathematics and programming problems, chess puzzles, and reasoning questions.
In active learning (AL), we focus on reducing the data annotation cost from the model training perspective. However, "testing", which often refers to the model