Reward Distance Comparisons Under Transition Sparsity
Nyanhongo, Clement, Henrique, Bruno Miranda, Santos, Eugene
–arXiv.org Artificial Intelligence
Reward comparisons are vital for evaluating differences in agent behaviors induced by a set of reward functions. Most conventional techniques utilize the input reward functions to learn optimized policies, which are then used to compare agent behaviors. However, learning these policies can be computationally expensive and can also raise safety concerns. Direct reward comparison techniques obviate policy learning but suffer from transition sparsity, where only a small subset of transitions are sampled due to data collection challenges and feasibility constraints. Existing state-of-the-art direct reward comparison methods are ill-suited for these sparse conditions since they require high transition coverage, where the majority of transitions from a given coverage distribution are sampled. When this requirement is not satisfied, a distribution mismatch between sampled and expected transitions can occur, leading to significant errors. This paper introduces the Sparsity Resilient Reward Distance (SRRD) pseudometric, designed to eliminate the need for high transition coverage by accommodating diverse sample distributions, which are common under transition sparsity. We provide theoretical justification for SRRD's robustness and conduct experiments to demonstrate its practical efficacy across multiple domains.
arXiv.org Artificial Intelligence
Apr-17-2025
- Genre:
- Research Report
- Experimental Study (0.92)
- New Finding (1.00)
- Research Report
- Industry:
- Government > Military (0.92)
- Health & Medicine > Health Care Providers & Services (0.67)
- Leisure & Entertainment > Games
- Computer Games (0.68)
- Technology:
- Information Technology > Artificial Intelligence
- Machine Learning
- Reinforcement Learning (1.00)
- Statistical Learning (1.00)
- Natural Language (1.00)
- Representation & Reasoning > Agents (1.00)
- Robots (1.00)
- Machine Learning
- Information Technology > Artificial Intelligence