Country
VisualAdversarialImitationLearning usingVariationalModels
Behaviour cloning (BC) is a classic algorithm to imitate expert demonstrations [7], which uses supervised learning to greedily match the expert behaviour at demonstrated expert states. Due to environmentstochasticity,covariateshift,andpolicyapproximationerror,theagentmaydriftaway from the expert state distribution and ultimately fail to mimic the demonstrator [8].
Appendixfor RiemannianContinuousNormalizingFlows
In the following, we provide a brief overview of Riemannian geometry and constant curvature manifolds, specifically the Poincarรฉ ball and the hypersphere models. Sphere In the two-dimensional settingd = 2, we rely on polar coordinates to parametrize the sphere S2. In the following subsection we remind that this regularization term can also be motivated from an estimator'svarianceperspective. 5 D.2 Frobeniusnorm Hutchinson'sestimator Hutchinson'sestimator(Hutchinson,1990)isasimple waytoobtain a stochastic estimate ofthetrace ofamatrix. The variance of this estimator thus depends on the Frobenius norm of the vector's field Jacobian Thenฮณ(tn) is also a Cauchy sequence by Equation 16. So for every sequence (tn) in (a,b) that converges tob, we have that(ฮณ(tn)) converges top.
1757af1fe1429801bdf3abf5600f8bba-Supplemental-Conference.pdf
The point of the highest Top1 accuracy is represented by five-pointed star for bothapproaches. In addition, as shown in Fig. r3, theoptimal balance coefficients forMobileNetV3 onCIFAR10, MobileNetV3 onCIFAR100 and ResNet50 on CIFAR100 are 5, 10 and 10 respectively. In the logits distillation, the main loss coefficient is 0.8 and distillation coefficient is 0.2. In the feature distillation, the feature distillation coefficient of each stage is 0.02. ForCIFAR100, we train all 4 networks utilizing standard training setting and do the same test above.