Catastrophic-risk-aware reinforcement learning with extreme-value-theory-based policy gradients

Open in new window