Goto

Collaborating Authors

 Deep Learning




aeb7b30ef1d024a76f21a1d40e30c302-Supplemental.pdf

Neural Information Processing Systems

In D, we show the proofs of the two propositions formulated in the main text. We further provide the results of evaluating our models using various metrics other than ECE (like AdaECE, Classwise-ECE, MCE and NLL). To empirically observe this, we use the ResNet-50 network used for the analysis in 3. We divide In Figure A.1, we also show Here we show why focal loss favours accurate but relatively less confident solutions. Experimentally, we found the solution of the cross-entropy and focal loss equations, i.e. the value of the predicted We consider a binary classification problem. Figures C.2 (b) and (c) show that running gradient descent with cross-entropy (CE) and focal loss (FL) both gives the same decision regions i.e. the weight vector Figure C.2: (a): Confidence of mis-classifications (b): Decision boundary of linear classifier trained Here we provide the proofs of both the propositions presented in the main text.


Calibrating Deep Neural Networks using Focal Loss

Neural Information Processing Systems

Deep Neural Networks (DNNs) makes their predictions hard to rely on. Ideally, we want networks to be accurate, calibrated and confident.







Supplementary Material: Reverse engineering recurrent neural networks with Jacobian switching linear dynamical systems

Neural Information Processing Systems

In general, we have found the JSLDS loss function strengths to be relatively easy to select (see example settings in the specific experiment sections below). RNN's fixed points or slow points would defeat the primary purpose of the method. However, other variations are possible. We set the number of timesteps T = 25. We trained both methods with the Adam optimizer with default settings.