Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling

Open in new window