Dataset distillation can be formulated as a bi-level meta-learning problem where the outer loop optimizes the metadataset and the inner loop trains a model on the distilled data.
Therefore, they seemtobegoodcandidatestobuild SSPsolutions, since Property 1 statesthattheidealsolverU is permutation-equivariant (thiswillbeconfirmedby Corollary 1).
The goal is to find the global optimal arm, and agents are able to pull any arm; however, they can only observe the reward when the selected arm is local.