An open vision-language model for diverse medical applications
20 hours ago
- MedGemma is a collection of medical vision-language foundation models based on Gemma 3, with variants of 4B, 27B, and 27B Text parameters.
- It outperforms similarly sized generative models on medical benchmarks, showing 2.6–10% improvement in medical image QA, 15.5–18.1% in chest X-ray classification, and 10.8% in agentic tasks.
- The models include MedSigLIP, a medical image encoder trained on millions of image-text pairs, enabling data-efficient classification and retrieval across modalities.
- Fine-tuning MedGemma provides robust task-specific adaptation, particularly effective with limited training data compared to base Gemma 3.
- MedGemma retains general-purpose capabilities while excelling in medical tasks, offering advantages in cost, local deployment, and data sovereignty.
- The open release of MedGemma and MedSigLIP aims to accelerate medical research and downstream AI applications in healthcare.