Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Open in new window