We address these flaws through a study of alternative self-supervised feature extractors, find that the semantic information encoded by individual networks strongly depends on their training procedure, and show that DINOv2-ViT -L/14 allows for much richer evaluation of generative models.
Further, we clarify that the term 'self-supervised' has different meanings in DA and in pretraining [S23, S24, S25, S26]. Problem definitions for both papers are different.
To address this challenge, methods need to accommodate target domains with different temporal dynamics and be capable of doing so without seeing any target examples during pre-training.