Goto

Collaborating Authors

 Technology



Appendix

Neural Information Processing Systems

Potential games are appealing as they provide guarantees about the existence of Nash equilibria without requiring optimizingoveracompact set. The learning rate isset to6 10 4 atinitialization and issequentially lowered during training by a factor of 0.35 every 80 training steps, in the same way for all experiments. Similar to [52], we normalize the initial dictionaryD0 by its largest singular value as explained in the main paper in Section 3.4. Divergence is monitored by computing the loss on the training set every 10 epochs.





Large-batchOptimizationforDenseVisualPredictions

Neural Information Processing Systems

At thet-th backward propagation step, we can derive the gradient il(wt)toupdatei-th module inM. The number in the bracket represents the batch size. We see that when the batch size is small (i.e.,32), the gradientvariancesaresimilar. N and K indicate the number of FPN levels and region proposals fed into the detection head. To evaluate this assumption, as shown in Figure 1, we have three observations. As illustrated by the second figure in Figure 1, the gradient misalignment phenomenon between detection head and backbone has been reduced.