Goto

Collaborating Authors

 Technology


Supplementary Material for Flat Seeking Bayesian Neural Networks Van-Anh Nguyen 1 Tung-Long Vuong

Neural Information Processing Systems

The proof can be found in Chapter 27 of [6]. For the non-flat version, the update is similar to the mini-batch SGD except that we add small Gaussian noises to the particle models. In Section 4.2 of the main paper, we provide a comprehensive analysis of the performance concerning In the experiments presented in Tables 1 and 2 in the main paper, we train all models for 300 epochs using SGD, with a learning rate of 0.1 and a cosine schedule. For the baseline of the Deep-Ensemble, SGLD, SGVB and SGVB-LRT methods, we reproduce results following the hyper-parameters and processes as our flat versions. ImageNet: This is a large and challenging dataset with 1000 classes.




Processing of missing data by neural networks

Neural Information Processing Systems

Our idea is to replace typical neuron's response in the firsthiddenlayerbyitsexpected value. Thisapproach canbeappliedforvarious types ofnetworksatminimal costintheirmodification. Moreover,incontrast to recent approaches, it does not require complete data for training. Experimental results performed ondifferent types ofarchitectures showthatourmethod gives better results than typical imputation strategies and other methods dedicated for incompletedata.






DoResidualNeuralNetworksdiscretizeNeural OrdinaryDifferentialEquations?

Neural Information Processing Systems

Neural ODEs also provide atheoretical framework to study deep learning models from the continuous viewpoint, using the arsenal of ODE theory [40, 25, 41]. Importantly, they can also be seen as the continuous analog of ResNets.