Goto

Collaborating Authors

 Statistical Learning





Appendix A Training details

Neural Information Processing Systems

Models are trained with Stochastic Gradient Descent with momentum equal to 0.9 [ We use a learning rate annealing scheme, decreasing the learning rate by a factor of 0.1 every 30 epochs. We train all models for 150 epochs. Then, we select the best learning rate and weight decay for each method and run 5 different seeds to report mean and standard deviation. We use the validation set of ImageNet to perform cross-validation and report performance on it. In section G we train the Augerino method on top of the Resnet-18 architecture.


AThe

Neural Information Processing Systems

B.2.1 Metrics Theevaluationmetricsweuseare Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and Mean Absolute Percentage Error (MAPE).



tandx

Neural Information Processing Systems

BytheMarkovian assumption forlatent state vectors, the Hessian matrix is tri-block diagonal. To facilitate convergence, we initialize the Newton update with a smoothing estimate bylocalGaussian approximation. Theforwardfiltering foradynamic Poisson modelhas been previously described (Eden etal., 2004), and we use anadditional backward pass tosmooth (Rauchetal.,1965). Without constraints, the sampling ofh(j), g(j) and ฯƒ2(j) is the same as shown previously. The update of A(j), b(j) and Q(j) is the standard multivariate Bayesian linear regression.


7b39f4512a2e3899edcc59c7501f3cd4-Paper-Conference.pdf

Neural Information Processing Systems

The LDS model is built on the state-space model and assumes latent factors evolvewith linear dynamics. Ontheother hand, GPFAmodels thelatent vectors by non-parametric Gaussian processes.