Somewhat surprisingly, learning in a Markov Decision Process is most often considered under the performance criteria ofconsistency or regret minimization(see e.g.
This formalizes a balance between learning low-dimensional representations and minimizing complexity/irregularity in the feature maps, allowing the network to learn the'right' inner dimension.