TimeDiscretization-Invariant SafeActionRepetitionforPolicyGradientMethods

Feb-7-2026, 06:52:57 GMT–Neural Information Processing Systems

In reinforcement learning, continuous time is often discretized by a time scale δ, to which the resulting performance is known to be highly sensitive. In this work, we seek tofind aδ-invariantalgorithm for policygradient (PG) methods, which performs well regardless of the value ofδ. We first identify the underlying reasons that cause PG methods to fail asδ 0, proving that the variance of the PG estimator can diverge to infinity in stochastic environments under a certain assumption of stochasticity. While durative actions or action repetition can be employed to haveδ-invariance, previous action repetition methods cannot immediately react to unexpected situations in stochastic environments. We thus propose a novelδ-invariant method namedSafe Action Repetition (SAR) applicable to any existing PG algorithm. SAR can handle the stochasticity of environments byadaptivelyreacting tochanges instates during action repetition.

artificial intelligence, machine learning, reinforcement learning, (18 more...)

Neural Information Processing Systems

Feb-7-2026, 06:52:57 GMT

Conferences PDF

Add feedback

Country:
- Europe > France (0.04)
- Asia
  - Middle East > Jordan (0.04)
  - Vietnam > Long An Province (0.04)

Technology:
- Information Technology > Artificial Intelligence > Machine Learning > Reinforcement Learning (0.69)

Duplicate Docs Excel Report

Title
024677efb8e4aee2eaeef17b54695bbe-Paper.pdf

Similar Docs Excel Report more

Title	Similarity	Source
None found