Goto

Collaborating Authors

 Country



ExplicablePolicySearch

Neural Information Processing Systems

The state space contains the position and velocity ofthe ego-vehicle and nearby vehicles. The action space consists offiveactions: accelerate, brake,idle, steer left, and steer right.


ExplicablePolicySearch

Neural Information Processing Systems

Human teammates often form conscious andsubconscious expectations ofeach other during interaction. Teaming success is contingent on whether such expectations can be met. Similarly,for an intelligent agent tooperate beside ahuman, it must consider the human's expectation of its behavior. Disregarding such expectations can lead to the loss of trust and degraded team performance. A key challenge here is that the human's expectation may not align with the agent's optimal behavior,e.g., duetothehuman'spartial orinaccurate understanding of thetaskdomain.




State Regularized Policy Optimization on Data with Dynamics Shift

Neural Information Processing Systems

We then demonstrate a lower-bound performance guarantee on policies regularized by the stationary state distribution. In practice, SRPO can be an add-on module to context-based algorithms in both online and offline RL settings.