Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
Guo, Jizhou, Wu, Zhaomin, Yang, Hanchen, Yu, Philip S.
Enhancing Large Language Model (LLM)'s performance with best-of-N sampling is effective and has attracted significant attention. However, it is computationally prohibitive due to massive, data-hungry text-based reward models. By changing the data source from text to hidden states, we introduce SWIFT (Simple Weighted Intrinsic Feedback Technique), a novel, lightweight technique that leverages the rich information embedded in LLM hidden states to address these issues, which operates on token-level and consists of only linear layers. Extensive experiments show that SWIFT outperforms baselines with less than 0.005% of the parameters of baselines, requiring only a few samples for training, demonstrating significant efficiency improvement. SWIFT's robust scalability, applicability to some closed-source models via logits, and ability to be combined with traditional reward models to yield further performance gains underscore its practical value.
Jul-30-2025
- Country:
- North America
- Canada (0.04)
- United States
- New Mexico > Bernalillo County
- Albuquerque (0.04)
- Illinois > Cook County
- Chicago (0.04)
- Florida > Miami-Dade County
- Miami (0.04)
- New Mexico > Bernalillo County
- Asia
- North America
- Genre:
- Research Report > New Finding (0.93)
- Technology: