Goto

Collaborating Authors

 Search


on ResNet-50 and by 7.3% on MobileNetV2

Neural Information Processing Systems

Our gains are indeed large. EvoNorm-S0 is the state-of-the-art in the small batch size regime (Table 4), outperforming BN-ReLU by 7.8% We achieve clear gains over other influential works such as GroupNorm (GN). We'd also like to emphasize that EvoNorms beat BN-ReLU on 12 (out of 14) different classification models/training These are significant considering the predominance of BN-ReLU in ML models. R3: "the overall search algorithm lacks some novelty." "yet another AutoML paper" (with the expectation that some fancy search algorithms must be proposed), but rather under R2, R4: Can EvoNorms generalize to deeper variants (e.g., ResNet-101) and architecture families not included MnasNet, EfficientNet-B5, Mask R-CNN + FPN/SpineNet and BigGAN-none of them was used during search.


Achieving Near-Optimal Convergence for Distributed Minimax Optimization with Adaptive Stepsizes

Neural Information Processing Systems

Sharma et al. (2022) provide Y ang et al. (2022a) integrate Local SGDA with stochastic gradient estimators to eliminate the More recently, Zhang et al. (2023) adopt compressed momentum methods with Local SGD to increase the communication efficiency of the algorithm. For centralized nonconvex minimax problems, Y ang et al. (2022b) show that, even in deterministic settings, GDA-based methods necessitate the timescale separation of the stepsizes for primal and dual updates.








A Proof of proposition

Neural Information Processing Systems

Let's assume we apply a random CCW torsion rotation of angle We detail here the formulae used in section section 2.4. Similar to AlphaFold [Senior et al., 2020], we fit distances using normal distributions and angles Such cases require a special treatment. So far, we haven't tackled the following difficulty: Examples are hydrogen groups as in Figure 1. We propose a new loss function based on eq. The EMD computation cannot be parallelized in mini-batches in the current version of the library, but everything else is batch-parallelizable in our model (e.g., The training stage happens without assembling the full conformer.