Goto

Collaborating Authors

 Technology








9b8b50fb590c590ffbf1295ce92258dc-AuthorFeedback.pdf

Neural Information Processing Systems

For example, when solving RL problems such as Atari7 games, we may test different representation methods. Fortheaveragereward30 setting, it is still an open question whether S-bounds areachievable. Ourapproach canbeadapted totheepisodic31 case when the regret bounds would benefit from the improved bounds available in this setting. The A-dependence is optimal as for UCRL2, while the optimal dependence onS is still an open question (also46 for the MDP case). The optimal dependence on|ฮฆ| in our setting is also open.



fe248e22b241ae5a9adf11493c8c12bc-Supplemental-Conference.pdf

Neural Information Processing Systems

In practice, however, the runtime is much smaller becauseofpruning. With this change in place, the solver can search for incomplete trees. KamPost differs from CART by using a different splitting criterion and by its post-relabelling of theleafnodes. The results also confirm thefindings from Figure 1that thevariance inthediscrimination value is often high, specifically for the small datasets. This means that for those instances it is difficult to generalize and overfitting interms ofdiscrimination isstill happening.