Visual Conceptual Blending with Large-scale Language and Vision Models

Jun-26-2021–arXiv.org Artificial Intelligence

We ask the question: to what extent can recent large-scale language and image generation models blend visual concepts? Given an arbitrary object, we identify a relevant object and generate a single-sentence description of the blend of the two using a language model. We then generate a visual depiction of the blend using a text-based image generation model. Quantitative and qualitative evaluations demonstrate the superiority of language models over classical methods for conceptual blending, and of recent large-scale image generation models over prior models for the visual depiction.

large-scale language and vision model, obj type annot subtype link, parent 231 0, (1 more...)

arXiv.org Artificial Intelligence

Jun-26-2021

arXiv.org Web Page

Add feedback

Genre:
- Research Report (0.40)

Technology:
- Information Technology > Artificial Intelligence > Vision (1.00)