Technology
4a1c2f4dcf2bf76b6b278ae40875d536-AuthorFeedback.pdf
We consider hk to be either Eq. 4 (proximal point), Eq. 9 (proximal gradient), or Eq. In the appendix, we conduct an experiment with an extremely7 high condition number (much higher than what is traditionally used for these problems). We agree that the clarity of our paper is subject to improvement and we thank the reviewer for his12 suggestions,whichwewilltakeintoaccount,ifthepaperisaccepted. NIPS 2014.", which displays moderate gains on text classification tasks. We agree that direct acceleration methods are appealing.
57d8ebf4c2f050a6485f370d47656a9e-Supplemental-Conference.pdf
In this section, we report the hyperparameters of each base model used in our paper, details in Table 2. The only hyperparameter that is tuned is done per dataset using a 10% validation split. In this Section, we discuss the experimental convergence of our U-DIF algorithm to the global optimum. In order to approximately compute the true global optimum, we use the following numerical scheme. (exact numbers vary by network and are given in Figure 4).