CogVLM: Visual Expert for Pretrained Language Models
–Neural Information Processing Systems
We introduce CogVLM, a powerful open-source visual language foundation model. Different from the popular \emph{shallow alignment} method which maps image features into the input space of language model, CogVLM bridges the gap between the frozen pretrained language model and image encoder by a trainable visual expert module in the attention and FFN layers. As a result, CogVLM enables a deep fusion of vision language features without sacrificing any performance on NLP tasks. Codes and checkpoints are available at Github.
Neural Information Processing Systems
May-27-2025, 19:11:38 GMT
- Technology: