6 hours ago
- EmbeddingGemma 2 is a lightweight multimodal embedding model built on the Gemma 4 architecture, released under Apache 2.0 license with 740 million parameters.
- It extends beyond text to unify code, images, video, and audio in a shared embedding space, enabling cross-modal search and retrieval on consumer hardware.
- Best-in-class for its size, achieving leading scores on benchmarks like MTEB Code and MAEB, often matching or outperforming larger models.
- Modular design allows for text-only workloads with as little as 270M parameters, plus optional vision (170M) and audio (300M) encoders.
- Storage-efficient via Matryoshka Representation Learning (MRL), enabling vector dimension reduction from 768 to as low as 128, reducing storage up to 6x.
- Optimized for on-device performance with quantization, requiring ~191MB RAM for text-only and ~567MB for full multimodal model on a Google Pixel 11 Pro.
- Extended 8K token context window supports up to 5.5 minutes of audio, 29 images, or 58 video frames locally.
- Achieves significant 9.92-point improvement on code performance (MTEB Code) over the previous version, and sets new quality-per-parameter standards across modalities.
- Enables data privacy and offline functionality via local embedding generation, and works with generative models like Gemma 4 for on-device RAG pipelines.
- Available for download on Hugging Face, Kaggle, and integrated with tools like LiteRT, MediaPipe, transformers.js, and various development frameworks.