From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities

Zhang, Wanpeng, Xie, Zilong, Feng, Yicheng, Li, Yijiang, Xing, Xingrun, Zheng, Sipeng, Lu, Zongqing

arXiv.org Artificial Intelligence 

Multimodal Large Language Models have made significant strides in integrating visual and textual information, yet they often struggle with effectively aligning these modalities. We introduce a novel image tokenizer that bridges this gap by applying the principle of Byte-Pair Encoding (BPE) to visual data. Unlike conventional approaches that rely on separate visual encoders, our method directly incorporates structural prior information into image tokens, mirroring the successful tokenization strategies used in text-only Large Language Models. This innovative approach enables Transformer models to more effectively learn and reason across modalities. Through theoretical analysis and extensive experiments, we demonstrate that our BPE Image Tokenizer significantly enhances MLLMs' multimodal understanding capabilities, even with limited training data. Our method not only improves performance across various benchmarks but also shows promising scalability, potentially paving the way for more efficient and capable multimodal foundation models. The development of Multimodal Large Language Models (MLLMs) has made significant progress (Yin et al., 2023; Team et al., 2023; Liu et al., 2024b). While this approach allows training data to align well with these modality-specific designs, it often struggles to achieve a unified understanding of multimodal information (Team, 2024). The primary reason for this limitation could be that while encoders of other modalities can learn rich information, without the assistance of the corresponding decoders, LLMs cannot fully comprehend the complex patterns contained within the embeddings provided by the encoder.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found