Technology
c20bb2d9a50d5ac1f713f8b34d9aac5a-AuthorFeedback.pdf
Both initialization methods ondownstream tasks canachievesimilar performance, butinitializing from BERT-base33 reduces the number of learning steps. In order to shorten the training time of our large-size model, we initialize34 it from BERT-large. We will also release a model trained from scratch. We further trained BERT-large using the35 same hyper-parameters, buttheresulted model didn'tsignificantly improvedownstream tasks compared tooriginal36 BERT-large.
Smoothed analysis of the low-rank approach for smooth semidefinite programs
Thomas Pumir, Samy Jelassi, Nicolas Boumal
Inprior work, ithas been shown that, when the constraints on the factorized variable regularly define a smooth manifold, providedk is large enough, for almost all cost matrices, all second-order stationary points (SOSPs) are optimal. Importantly, in practice, one can only compute points which approximately satisfy necessary optimality conditions, leading tothequestion: aresuch points also approximately optimal?