Compression Strategies for Efficient Multimodal LLMs in Medical Contexts

Khan, Tanvir A., Saha, Aranya, Swapnil, Ismam N., Haque, Mohammad A.

arXiv.org Artificial Intelligence 

We present a unified and efficient pipeline for deploying multimodal large language models (MLLMs) in domain-specific settings, with a focus on dermatological visual question answering (VQA). Real-time clinical workflows require low-latency models, yet the substantial memory and computational demands of LLMs and MLLMs hinder their deployment in resource-constrained edge or cloud environments. To tackle these challenges, we introduce a compression pipeline that integrates structural pruning, supervised fine-tuning (SFT), and quantization for task specific purposes. Our approach begins with structured pruning to remove redundant parameters, reducing the model's size while preserving essential functionality. We employ a novel layer pruning criterion for domain specific applications.