GLAMI-1M: A Multilingual Image-Text Fashion Dataset
Kosar, Vaclav, Hoskovec, Antonín, Šulc, Milan, Bartyzal, Radek
–arXiv.org Artificial Intelligence
We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 classes has high-quality annotations: all 100k images in the test set and 75% of the 1M training set were human-labeled. The paper presents baselines for image-text classification showing that the dataset presents a challenging fine-grained classification problem: The best scoring EmbraceNet model using both visual and textual features achieves 69.7% accuracy. Experiments with a modified Imagen model show the dataset is also suitable for image generation conditioned on text. The dataset, source code and model checkpoints are published at https://github.com/glami/glami-1m
arXiv.org Artificial Intelligence
Nov-17-2022
- Country:
- Oceania > Palau (0.04)
- North America
- Europe
- Czechia > Prague (0.05)
- Slovenia (0.04)
- Lithuania (0.04)
- Bulgaria (0.04)
- Spain (0.04)
- Slovakia (0.04)
- Greece (0.04)
- Latvia (0.04)
- Estonia (0.04)
- Romania (0.04)
- Croatia (0.04)
- Hungary (0.04)
- Poland (0.04)
- France > Provence-Alpes-Côte d'Azur
- Bouches-du-Rhône > Marseille (0.04)
- Denmark > Capital Region
- Copenhagen (0.04)
- Asia
- Middle East > Republic of Türkiye (0.04)
- Japan > Honshū
- Tōhoku > Fukushima Prefecture > Fukushima (0.04)
- Genre:
- Research Report (0.52)
- Technology: