Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine

Open in new window