Country
RandomShufflingBeatsSGDOnlyAfterMany EpochsonIll-ConditionedProblems
However, known lower bounds ignore the problem's geometry,including itscondition number,whereas theupper bounds explicitly depend on it. Perhaps surprisingly, we prove that when the condition number is taken into account, without-replacement SGDdoesnotsignificantly improveon withreplacement SGD in terms of worst-case bounds, unless the number of epochs (passes overthedata) islargerthanthecondition number.
LearningandTransferringSparseContextualBigrams withLinearTransformers
Weshowthat when trained from scratch,thetraining process can be split into an initial sample-intensive stage where the correlation is boosted from zero to a nontrivial value, followed by a more sample-efficient stageoffurther improvement. Additionally,weprovethat, provided anontrivial correlation between the downstream and pretraining tasks, finetuning from a pretrained model allowsustobypass the initial sample-intensivestage.