Goto

Collaborating Authors

 Statistical Learning


A Theoretical Analysis

Neural Information Processing Systems

Note 3: Consider a balanced continual learning dataset (e.g., Split-CIFAR100, Split-Mini-ImageNet) Note 4: Consider general continual learning datasets. The hyperparameter settings are summarized in Table 4. All models are optimized using vanilla SGD. For all experiments, we use the learning rate of 0.1 following the same setting as in Aljundi et al. Mai et al. reported (2021) considerable and consistent performance gains when replacing the Softmax classifier with the NCM classifier.


A Bayesian Inference over Neural Networks On a supervised model parameterized by W, we seek to infer the conditional distribution W | D

Neural Information Processing Systems

The prior and likelihood are both modelling choices. A.1 Likelihoods for BNNs The likelihood is purely a function of the model prediction ฮฆ As exact posterior inference via (11) is intractable, we instead rely on approximate inference algorithms, which can be broadly grouped into two classes based on their method of approximation. A concrete label can be obtained by choosing the class with highest output value. The Gaussian variational family is a common choice. Estimators for the integral in (15) are necessary.