The Impact of Quantization and Pruning on Deep Reinforcement Learning Models
Lu, Heng, Alemi, Mehdi, Rawassizadeh, Reza
–arXiv.org Artificial Intelligence
Reinforcement learning has been applied in many fields, including robotics, video games, and recently Reinforcement Learning with Human Feedback (RLHF) [1, 2, 3, 4, 5], has become common in large language models. RLHF methods mitigate biases inherent in language models themselves [4, 6, 7]. Reinforcement learning models that address real-world problems predominantly utilize continuous models based on neural network architecture, known as Deep Reinforcement Learning (DRL). DRL methods typically involve a world model, agents interacting with the world, and a reward function that evaluates the effectiveness of actions based on the agent's policy towards predefined objectives [8]. Depending on whether the algorithm learns a specific world model, DRL algorithms are categorized into model-based DRL algorithms and model-free DRL algorithms [9]. Model-free DRL methods generally fall into three main categories: deep Q-learning methods [10, 11, 12], policy gradient methods [13, 14, 15], and actor-critic methods [16, 17, 18, 19]. Unlike model-based algorithms, model-free approaches circumvent model bias and offer greater generalizability, which contributes to their popularity in RLHF applications [4, 20]. Neural networks, which are the backbone of DRL methods, are associated with high computational costs and, therefore, resource intensive.
arXiv.org Artificial Intelligence
Jul-5-2024
- Country:
- North America
- United States
- Massachusetts
- Suffolk County > Boston (0.04)
- Middlesex County > Natick (0.04)
- Florida > Broward County
- Fort Lauderdale (0.04)
- Massachusetts
- Puerto Rico > San Juan
- San Juan (0.04)
- United States
- Europe
- Asia > Middle East
- Jordan (0.04)
- North America
- Genre:
- Research Report (0.82)
- Industry:
- Information Technology (0.68)
- Leisure & Entertainment > Games
- Computer Games (0.34)
- Technology: