CogVLM: Visual Expert for Pretrained Language Models
–Neural Information Processing Systems
We introduce CogVLM, a powerful open-source visual language foundation model. Different from the popular \emph{shallow alignment} method which maps image features into the input space of language model, CogVLM bridges the gap between the frozen pretrained language model and image encoder by a trainable visual expert module in the attention and FFN layers. As a result, CogVLM enables a deep fusion of vision language features without sacrificing any performance on NLP tasks.
Neural Information Processing Systems
Mar-22-2026, 16:38:48 GMT
- Technology: