Deep Learning
Appendix: On the Overlooked Structure of Stochastic Gradients
Avila is a non-image dataset. A.3 Image classification on MNIST We perform the common per-pixel zero-mean unit-variance normalization as data preprocessing for MNIST. Pretraining Hyperparameter Settings: We train neural networks for 50 epochs on MNIST for obtaining pretrained models. The batch size is set to 1 and no weight decay is used, unless we specify them otherwise. As for other optimizer hyperparameters, we apply the default settings directly.
Supplementary Material
We use the PyTorch framework for our experiments. Similar to TD3, we implement our GRU-ODE in SAC. In this ablation study, we ask two questions in relation to numerical integration. Thus, simple numerical solvers are enough. We evaluate the time costs of different baselines on Walker-P environments.