Bayes' rule naturally allows for inference refinement in a streaming fashion, without the need to recompute posteriors from scratch whenever new data arrives.
This has motivated recent works to explore the benefits of incorporating publicly available data into private training, e.g., by pretraining a model on public data and then finetuning it using private data.
We show a fast-rate regret bound for IERM that allows for misspecified model classes and flexible choices of the optimization estimate, and we develop computationally tractable surrogate losses.
Code is available at https://github.com/xbyym/DLSR. If the test data does not follow the training distribution, the model could unintentionally produce nonsensical predictions, resulting in some misleading conclusions.
Lipschitz gradient, without needing to know neither the corresponding Lipschitz constants, nor the oracle's variance but enjoying the rates which are characteristic for algorithms which have the