Differentially Private Model-Based Offline Reinforcement Learning
Rio, Alexandre, Barlier, Merwan, Colin, Igor, Thomas, Albert
–arXiv.org Artificial Intelligence
We address offline reinforcement learning with privacy guarantees, where the goal is to train a policy that is differentially private with respect to individual trajectories in the dataset. To achieve this, we introduce DP-MORL, an MBRL algorithm coming with differential privacy guarantees. A private model of the environment is first learned from offline data using DP-FedAvg, a training method for neural networks that provides differential privacy guarantees at the trajectory level. Then, we use model-based policy optimization to derive a policy from the (penalized) private model, without any further interaction with the system or access to the input data. We empirically show that DP-MORL enables the training of private RL agents from offline data and we furthermore outline the price of privacy in this setting.
arXiv.org Artificial Intelligence
Feb-8-2024
- Country:
- North America > United States (0.04)
- Europe
- Portugal (0.04)
- France > Auvergne-Rhône-Alpes
- Finland > Northern Savo
- Kuopio (0.04)
- Asia > China
- Genre:
- Research Report (0.65)
- Industry:
- Information Technology > Security & Privacy (1.00)
- Technology: