Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models

Lian, Chenyu, Zhou, Hong-Yu, Liang, Dongyun, Qin, Jing, Wang, Liansheng

arXiv.org Artificial Intelligence 

-- Medical vision-language alignment through cross-modal contrastive learning shows promising performance in image-text matching tasks, such as retrieval and zero-shot classification. However, conventional cross-modal contrastive learning (CLIP-based) methods suffer from suboptimal visual representation capabilities, which also limits their effectiveness in vision-language alignment. In contrast, although the models pretrained via multimodal masked modeling struggle with direct cross-modal matching, they excel in visual representation. To address this contradiction, we propose ALTA (ALign Through Adapting), an efficient medical vision-language alignment method that utilizes only about 8% of the trainable parameters and less than 1/5 of the computational consumption required for masked record modeling. ALTA achieves superior performance in vision-language matching tasks like retrieval and zero-shot classification by adapting the pretrained vision model from masked record modeling. Additionally, we integrate temporal-multiview radiograph inputs to enhance the information consistency between radiographs and their corresponding descriptions in reports, further improving the vision-language alignment. Experimental evaluations show that ALTA outperforms the best-performing counterpart by over 4% absolute points in text-to-image accuracy and approximately 6% absolute points in image-to-text re-This work was supported by National Natural Science Foundation of China (Grant No. 62371409), Fujian Provincial Natural Science Foundation of China (Grant No. 2023J01005), Innovation and T ech-nology Fund under Hong Kong Innovation and T echnology Commission (Project No. ITS/202/23), Collaborative Research with World-leading Research Groups in The Hong Kong Polytechnic University (Project No. G-SACF), and a Shenzhen-Hong Kong-Macao Science and T echnology Plan Project (Category C Project) under Shenzhen Municipal Science and T echnology Innovation Commission (project no SGDX20230821092359002). Hong-Y u Zhou is with the Department of Biomedical Informatics, Harvard Medical School, Boston, USA (e-mail: whuzhouhongyu@gmail.com). Dongyun Liang is with the Department of Radiology, Zhongshan Hospital (Xiamen), Fudan University, Xiamen Municipal Clinical Research Center for Medical Imaging, Fujian Province Key Clinical Specialty for Medical Imaging, Xiamen Key Laboratory of Clinical T ransformation of Imaging Big Data and Artificial Intelligence, Xiamen 361015, China (email: ldy7372@163.com). Liansheng Wang is with the National Institute for Data Science in Health and Medicine, and the Department of Computer Science, School of Informatics, Xiamen University, Xiamen 361005, China (email: lswang@xmu.edu.cn). Jing Qin is with the Center for Smart Health, School of Nursing, The Hong Kong Polytechnic University, Hong Kong, China (e-mail: harry .qin@polyu.edu.hk).

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found