CREMA: Multimodal Compositional Video Reasoning via Efficient Modular Adaptation and Fusion

Open in new window