PixelBytes: Catching Unified Representation for Multimodal Generation

Furfaro, Fabien

arXiv.org Artificial Intelligence 

Recent advancements in artificial intelligence have led to increasingly generalist models, not by combining multiple specialized components (like Gato from DeepMind [28]), but by assigning simple tasks to models where emergent properties--complex behaviors arising from simpler underlying rules--appear. This is exemplified by generative language models such as GPT [7]. However, these models are constrained by their focus on language alone, failing to capture the full complexity of multimodal understanding [15]. To address this limitation, researchers have explored integrating Large Language Models (LLMs) with other modalities [24]. However, this approach often results in specialized model combinations without fostering new emergent properties. We propose "PixelBytes", a novel approach enabling unified training across modalities by representing diverse inputs in a single, cohesive format. Multimodal sequence generation, which involves creating coherent outputs combining various data types such as text, images, and numerical sequences, presents a significant challenge in artificial intelligence [2].