Wefurther propose newtraining methods todisentangle the embeddings, making them both distinctive signatures of the environments and tasks and effective building blocks for composing the policies.
In the reinforcement learning context, anOption means a temporally extended sequence of actions [30],andisregarded asuseful formanypurposes, such asspeeding uplearning, transferring skills across domains, and solving long-term planning problems.
A naiveimplementation of this approach leads to the dynamic component taking over the static one as the representation of the former is inherently more general and prone to overfitting.