Country
BeyondAesthetics: CulturalCompetencein Text-to-ImageModels
In particular, we apply this approach to build CUBE (CUltural BEnchmark forText-to-Image models), afirst-of-its-kind benchmark toevaluate cultural competence of T2I models.2 CUBE covers cultural artifacts associated with 8 countries across different geo-cultural regions and along 3 concepts: cuisine, landmarks, and art. CUBE consists of 1) CUBE-1K, a set of high-quality prompts thatenable theevaluation ofcultural awareness, and2)CUBE-CSpace, a larger dataset of cultural artifacts that serves as grounding to evaluate cultural diversity.
SupplementaryMaterialsforExemplarVAE: LinkingGenerativeModels,NearestNeighbor Retrieval,andDataAugmentation
This metric computes the variance of the mean of the latent encoding of the data points in each dimension of the latent space,Var(µφ(x)i), wherexis sampledfromthedataset. For hierarchical architectures the reported number is for thez2 which is the highest stochasticlayer. Toregularize the Exemplar VAE, we used leave-one-out and exemplar sub-sampling. That is why did not compare directly against a mixture model prior in the primary experimental section. Three different architectures are used in the experiments, described below.