Exploiting LMM-based knowledge for image classification tasks

Tzelepi, Maria, Mezaris, Vasileios

arXiv.org Artificial Intelligence 

Large Language Models (LLMs) [1, 2], such as GPT-3 [3] and GPT-4 [4], trained on vast amounts of data, have demonstrated exceptional performance in several downstream tasks over the recent few years, placing them squarely at the center of the research activity on Natural Language Processing (NLP) [5] and computer vision [6]. Considering vision recognition downstream tasks, in particular, the emergence of Vision-Language Models (VLMs), such as BLIP-2 [7] and CLIP [8], allowed for connecting image-based vision models with LLMs, forming Large Multimodal Models (LMMs) (also known as Multimodal Large Language Models) [9, 10]. For example, MiniGPT-4 [11] aligns a frozen visual encoder with a frozen LLM using a single projection layer. Besides, it is noteworthy that CLIP, which apart from achieving remarkable zero-shot performance on various downstream tasks through prompting, is a powerful feature extractor, has been extensively used in the recent literature for various applications [12, 13, 14]. In this work, our goal is to address image classification tasks (i.e., tasks of assigning a class label to an image based on its visual content), harnessing the emerging technology of LLMs/LMMs, in order to achieve improved performance in terms of classification accuracy. To achieve this goal, we pursue the direction of utilizing CLIP, proposing to further incorporate knowledge encoded in powerful foundation models such as MiniGPT-4. More specifically, CLIP is commonly utilized as a feature extractor for fitting a linear classifier on the extracted image embeddings and evaluating the performance on various datasets. In this paper, we propose to use MiniGPT-4 for obtaining semantic textual descriptions for each sample of the considered dataset, and then to use the extracted descriptions for feeding them to the textual encoder of CLIP and obtaining the corresponding text embeddings.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found