PixelBytes: Catching Unified Representation for Multimodal Generation
–arXiv.org Artificial Intelligence
Recent advancements in artificial intelligence have led to increasingly generalist models, not by combining multiple specialized components (like Gato from DeepMind [28]), but by assigning simple tasks to models where emergent properties--complex behaviors arising from simpler underlying rules--appear. This is exemplified by generative language models such as GPT [7]. However, these models are constrained by their focus on language alone, failing to capture the full complexity of multimodal understanding [15]. To address this limitation, researchers have explored integrating Large Language Models (LLMs) with other modalities [24]. However, this approach often results in specialized model combinations without fostering new emergent properties. We propose "PixelBytes", a novel approach enabling unified training across modalities by representing diverse inputs in a single, cohesive format. Multimodal sequence generation, which involves creating coherent outputs combining various data types such as text, images, and numerical sequences, presents a significant challenge in artificial intelligence [2].
arXiv.org Artificial Intelligence
Oct-20-2024